{
  "id": 544342,
  "title": "Is The AutoEncoder Can Be The winner Model:",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/544342",
  "author_name": "",
  "post_date": "2024-11-04T16:15:42.487233900Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>when i revise the public model i relize that the some highest scores Notebooks are using some architecture that build derivadede from autoencoder :<br>\n<a href=\"https://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528\" target=\"_blank\">https://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528</a>   #0.492<br>\nmy notebook: <a href=\"https://www.kaggle.com/code/saidkoussi/vae-pytorch-lightgbm-wih-cuda-df\" target=\"_blank\">https://www.kaggle.com/code/saidkoussi/vae-pytorch-lightgbm-wih-cuda-df</a>   #0.490</p>\n<p>what do you guess about the winning solutions?</p>",
  "messages": [
    {
      "id": "3036508",
      "postDate": "11/04/2024 16:15:42",
      "content": "<p>when i revise the public model i relize that the some highest scores Notebooks are using some architecture that build derivadede from autoencoder :<br>\n<a href=\"https://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528\" target=\"_blank\">https://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528</a>   #0.492<br>\nmy notebook: <a href=\"https://www.kaggle.com/code/saidkoussi/vae-pytorch-lightgbm-wih-cuda-df\" target=\"_blank\">https://www.kaggle.com/code/saidkoussi/vae-pytorch-lightgbm-wih-cuda-df</a>   #0.490</p>\n<p>what do you guess about the winning solutions?</p>",
      "rawMarkdown": "when i revise the public model i relize that the some highest scores Notebooks are using some architecture that build derivadede from autoencoder :\nhttps://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528   #0.492\nmy notebook: https://www.kaggle.com/code/saidkoussi/vae-pytorch-lightgbm-wih-cuda-df   #0.490\n\nwhat do you guess about the winning solutions?",
      "votes": null
    },
    {
      "id": "3036732",
      "postDate": "11/04/2024 21:33:42",
      "content": "<p>Good question. It seems your final submission is a voting of submission 1,2,3, and only submission1 is using an autoencoder + KNN impute + feature engineering, and only time-series is autoencoded.<br>\nHave you tried using only submission1 and what's the public score?<br>\nYou can't tell unless you try doing the above mentioned one at a time.<br>\nI'm interested to know as well.<br>\nGood luck.</p>",
      "rawMarkdown": "Good question. It seems your final submission is a voting of submission 1,2,3, and only submission1 is using an autoencoder + KNN impute + feature engineering, and only time-series is autoencoded.\nHave you tried using only submission1 and what's the public score?\nYou can't tell unless you try doing the above mentioned one at a time.\nI'm interested to know as well.\nGood luck.",
      "votes": null
    },
    {
      "id": "3036963",
      "postDate": "11/05/2024 06:04:10",
      "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> i tried to ensemble various tree-based models with various autoencoder architecture , but the scores are unstable, so we must be carefull about data leckage, focusing on just increasing leadboard can be disapoints in the private leadboard , i think we must build robust model that has high cv , leadboard score </p>",
      "rawMarkdown": "tomyuen i tried to ensemble various tree-based models with various autoencoder architecture , but the scores are unstable, so we must be carefull about data leckage, focusing on just increasing leadboard can be disapoints in the private leadboard , i think we must build robust model that has high cv , leadboard score",
      "votes": null
    },
    {
      "id": "3037280",
      "postDate": "11/05/2024 14:58:18",
      "content": "<p>I can tell you that the autoencoder is not correctly applied in these notebooks because it comes from the same copy-paste source where the original author made a mistake in applying it on data - so take it like a hint for debugging.</p>",
      "rawMarkdown": "I can tell you that the autoencoder is not correctly applied in these notebooks because it comes from the same copy-paste source where the original author made a mistake in applying it on data - so take it like a hint for debugging.",
      "votes": null
    },
    {
      "id": "3037289",
      "postDate": "11/05/2024 15:09:03",
      "content": "<p>I also used a voting strategy and tried to submit each submission separately. </p>\n<p>Submitting separately leads to worse results than using a voting strategy for the final submission. Maybe because voting between submissions reduces model bias which leads to better generalization, single models might vary widely in predictions for certain challenging cases.</p>\n<p>I saw great improvements of my LB scores since I started voting between submissions.</p>\n<p>When using voting strategies we also need to consider that voting can mask the underlying variability of each model’s performance on different data samples.</p>\n<p>In other words, the ensemble might look consistent on the public set, but individual models may vary widely on certain data patterns, indicating that the model ensemble’s true generalization ability may be weaker than it appears.</p>\n<p>This can have an impact in our case : Kaggle test splits (public vs. private) are typically designed to represent the distribution of the entire test set. However, there can still be subtle distributional shifts between the public and private splits due to random sampling.</p>\n<p>If the ensemble voting strategy unintentionally overfits to specific quirks in the public test data, it may generalize poorly on the private test set. The final score on the private leaderboard could then be lower than expected.</p>\n<p>To briefly conclude on this, voting currently leads to better leaderboard scores, but it doesn’t mean those scores will translate onto the final test set and leaderboard. </p>",
      "rawMarkdown": "I also used a voting strategy and tried to submit each submission separately. \n\nSubmitting separately leads to worse results than using a voting strategy for the final submission. Maybe because voting between submissions reduces model bias which leads to better generalization, single models might vary widely in predictions for certain challenging cases.\n\nI saw great improvements of my LB scores since I started voting between submissions.\n\nWhen using voting strategies we also need to consider that voting can mask the underlying variability of each model’s performance on different data samples.\n\nIn other words, the ensemble might look consistent on the public set, but individual models may vary widely on certain data patterns, indicating that the model ensemble’s true generalization ability may be weaker than it appears.\n\nThis can have an impact in our case : Kaggle test splits (public vs. private) are typically designed to represent the distribution of the entire test set. However, there can still be subtle distributional shifts between the public and private splits due to random sampling.\n\nIf the ensemble voting strategy unintentionally overfits to specific quirks in the public test data, it may generalize poorly on the private test set. The final score on the private leaderboard could then be lower than expected.\n\nTo briefly conclude on this, voting currently leads to better leaderboard scores, but it doesn’t mean those scores will translate onto the final test set and leaderboard.",
      "votes": null
    },
    {
      "id": "3037309",
      "postDate": "11/05/2024 15:24:08",
      "content": "<p>VotingRegressor doesn't mask anything - it returns the average of the results, it's just a fancy np.mean(your_results).<br>\nHere is the official description - \"A voting regressor is an ensemble meta-estimator that fits several base regressors, each on the whole dataset. Then it averages the individual predictions to form a final prediction.\"</p>",
      "rawMarkdown": "VotingRegressor doesn't mask anything - it returns the average of the results, it's just a fancy np.mean(your_results).\nHere is the official description - \"A voting regressor is an ensemble meta-estimator that fits several base regressors, each on the whole dataset. Then it averages the individual predictions to form a final prediction.\"",
      "votes": null
    },
    {
      "id": "3037321",
      "postDate": "11/05/2024 15:39:43",
      "content": "<p>You’re absolutely right that VotingRegressor performs a straightforward averaging of predictions and doesn’t inherently mask variability. </p>\n<p>However, in this case, we’re not talking about VotingRegressors. Voting between submissions uses majority voting, which is different from averaging.</p>\n<p>By using majority voting rather than averaging, the aim is to reduce the impact of outlier predictions from individual models. This approach is particularly helpful for discrete ordinal predictions, where taking the mode can reduce the influence of individual model biases, as it discards predictions that diverge too much from the majority. This is less likely to happen with simple averaging. </p>\n<p>I wish you good luck in this competition. </p>",
      "rawMarkdown": "You’re absolutely right that VotingRegressor performs a straightforward averaging of predictions and doesn’t inherently mask variability. \n\nHowever, in this case, we’re not talking about VotingRegressors. Voting between submissions uses majority voting, which is different from averaging.\n\nBy using majority voting rather than averaging, the aim is to reduce the impact of outlier predictions from individual models. This approach is particularly helpful for discrete ordinal predictions, where taking the mode can reduce the influence of individual model biases, as it discards predictions that diverge too much from the majority. This is less likely to happen with simple averaging. \n\nI wish you good luck in this competition.",
      "votes": null
    },
    {
      "id": "3037335",
      "postDate": "11/05/2024 15:48:03",
      "content": "<p>I think not always the autocoder can be the winner model.They might be the \"winner\" for tasks that rely on dimensionality reduction, feature extraction, or anomaly detection, especially when data is unlabeled.</p>",
      "rawMarkdown": "I think not always the autocoder can be the winner model.They might be the \"winner\" for tasks that rely on dimensionality reduction, feature extraction, or anomaly detection, especially when data is unlabeled.",
      "votes": null
    },
    {
      "id": "3037591",
      "postDate": "11/05/2024 20:57:10",
      "content": "<p>I frequently observed the <code>GaussianNoise</code> layer receiving input data at the start of the AutoEncoder architectures, but here it's skipped, perhaps for some reason. On the step of scaling, I also noticed that the Gauss Rank Scaler was used rather than other scalers. Another addition is a Swap Noise. These three things I associate with Denoising AutoEncoder primarily. Here, as I can follow, AutoEncoder does feature extraction, so noise is not necessary, although it may be useful as a form of regularization to learn a more robust representation and improve generalization since there is a lack of data, but the issue of unidentified randomness still remains, thus how to test anything is a question.</p>",
      "rawMarkdown": "I frequently observed the `GaussianNoise` layer receiving input data at the start of the AutoEncoder architectures, but here it's skipped, perhaps for some reason. On the step of scaling, I also noticed that the Gauss Rank Scaler was used rather than other scalers. Another addition is a Swap Noise. These three things I associate with Denoising AutoEncoder primarily. Here, as I can follow, AutoEncoder does feature extraction, so noise is not necessary, although it may be useful as a form of regularization to learn a more robust representation and improve generalization since there is a lack of data, but the issue of unidentified randomness still remains, thus how to test anything is a question.",
      "votes": null
    },
    {
      "id": "3037690",
      "postDate": "11/06/2024 02:27:03",
      "content": "<p><a href=\"https://www.kaggle.com/sumit08\" target=\"_blank\">@sumit08</a> yes agree bout some portion of the data are unabled , so autoencoder may have the upper hand</p>",
      "rawMarkdown": "sumit08 yes agree bout some portion of the data are unabled , so autoencoder may have the upper hand",
      "votes": null
    },
    {
      "id": "3041461",
      "postDate": "11/10/2024 11:47:24",
      "content": "<p>Can you explain what is the mistake? thank you.</p>",
      "rawMarkdown": "Can you explain what is the mistake? thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3036732,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "11/04/2024 21:33:42",
      "content": "<p>Good question. It seems your final submission is a voting of submission 1,2,3, and only submission1 is using an autoencoder + KNN impute + feature engineering, and only time-series is autoencoded.<br>\nHave you tried using only submission1 and what's the public score?<br>\nYou can't tell unless you try doing the above mentioned one at a time.<br>\nI'm interested to know as well.<br>\nGood luck.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3036963,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "11/05/2024 06:04:10",
          "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> i tried to ensemble various tree-based models with various autoencoder architecture , but the scores are unstable, so we must be carefull about data leckage, focusing on just increasing leadboard can be disapoints in the private leadboard , i think we must build robust model that has high cv , leadboard score </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3037289,
          "author_name": "alixnguyen",
          "author_url": "",
          "post_date": "11/05/2024 15:09:03",
          "content": "<p>I also used a voting strategy and tried to submit each submission separately. </p>\n<p>Submitting separately leads to worse results than using a voting strategy for the final submission. Maybe because voting between submissions reduces model bias which leads to better generalization, single models might vary widely in predictions for certain challenging cases.</p>\n<p>I saw great improvements of my LB scores since I started voting between submissions.</p>\n<p>When using voting strategies we also need to consider that voting can mask the underlying variability of each model’s performance on different data samples.</p>\n<p>In other words, the ensemble might look consistent on the public set, but individual models may vary widely on certain data patterns, indicating that the model ensemble’s true generalization ability may be weaker than it appears.</p>\n<p>This can have an impact in our case : Kaggle test splits (public vs. private) are typically designed to represent the distribution of the entire test set. However, there can still be subtle distributional shifts between the public and private splits due to random sampling.</p>\n<p>If the ensemble voting strategy unintentionally overfits to specific quirks in the public test data, it may generalize poorly on the private test set. The final score on the private leaderboard could then be lower than expected.</p>\n<p>To briefly conclude on this, voting currently leads to better leaderboard scores, but it doesn’t mean those scores will translate onto the final test set and leaderboard. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3037309,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "11/05/2024 15:24:08",
              "content": "<p>VotingRegressor doesn't mask anything - it returns the average of the results, it's just a fancy np.mean(your_results).<br>\nHere is the official description - \"A voting regressor is an ensemble meta-estimator that fits several base regressors, each on the whole dataset. Then it averages the individual predictions to form a final prediction.\"</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3037321,
                  "author_name": "alixnguyen",
                  "author_url": "",
                  "post_date": "11/05/2024 15:39:43",
                  "content": "<p>You’re absolutely right that VotingRegressor performs a straightforward averaging of predictions and doesn’t inherently mask variability. </p>\n<p>However, in this case, we’re not talking about VotingRegressors. Voting between submissions uses majority voting, which is different from averaging.</p>\n<p>By using majority voting rather than averaging, the aim is to reduce the impact of outlier predictions from individual models. This approach is particularly helpful for discrete ordinal predictions, where taking the mode can reduce the influence of individual model biases, as it discards predictions that diverge too much from the majority. This is less likely to happen with simple averaging. </p>\n<p>I wish you good luck in this competition. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3037280,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "11/05/2024 14:58:18",
      "content": "<p>I can tell you that the autoencoder is not correctly applied in these notebooks because it comes from the same copy-paste source where the original author made a mistake in applying it on data - so take it like a hint for debugging.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3041461,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "11/10/2024 11:47:24",
          "content": "<p>Can you explain what is the mistake? thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3037335,
      "author_name": "sumit08",
      "author_url": "",
      "post_date": "11/05/2024 15:48:03",
      "content": "<p>I think not always the autocoder can be the winner model.They might be the \"winner\" for tasks that rely on dimensionality reduction, feature extraction, or anomaly detection, especially when data is unlabeled.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3037690,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "11/06/2024 02:27:03",
          "content": "<p><a href=\"https://www.kaggle.com/sumit08\" target=\"_blank\">@sumit08</a> yes agree bout some portion of the data are unabled , so autoencoder may have the upper hand</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3037591,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "11/05/2024 20:57:10",
      "content": "<p>I frequently observed the <code>GaussianNoise</code> layer receiving input data at the start of the AutoEncoder architectures, but here it's skipped, perhaps for some reason. On the step of scaling, I also noticed that the Gauss Rank Scaler was used rather than other scalers. Another addition is a Swap Noise. These three things I associate with Denoising AutoEncoder primarily. Here, as I can follow, AutoEncoder does feature extraction, so noise is not necessary, although it may be useful as a form of regularization to learn a more robust representation and improve generalization since there is a lack of data, but the issue of unidentified randomness still remains, thus how to test anything is a question.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3036508": "when i revise the public model i relize that the some highest scores Notebooks are using some architecture that build derivadede from autoencoder :\nhttps://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528   #0.492\nmy notebook: https://www.kaggle.com/code/saidkoussi/vae-pytorch-lightgbm-wih-cuda-df   #0.490\n\nwhat do you guess about the winning solutions?",
    "3036732": "Good question. It seems your final submission is a voting of submission 1,2,3, and only submission1 is using an autoencoder + KNN impute + feature engineering, and only time-series is autoencoded.\nHave you tried using only submission1 and what's the public score?\nYou can't tell unless you try doing the above mentioned one at a time.\nI'm interested to know as well.\nGood luck.",
    "3036963": "tomyuen i tried to ensemble various tree-based models with various autoencoder architecture , but the scores are unstable, so we must be carefull about data leckage, focusing on just increasing leadboard can be disapoints in the private leadboard , i think we must build robust model that has high cv , leadboard score",
    "3037280": "I can tell you that the autoencoder is not correctly applied in these notebooks because it comes from the same copy-paste source where the original author made a mistake in applying it on data - so take it like a hint for debugging.",
    "3037289": "I also used a voting strategy and tried to submit each submission separately. \n\nSubmitting separately leads to worse results than using a voting strategy for the final submission. Maybe because voting between submissions reduces model bias which leads to better generalization, single models might vary widely in predictions for certain challenging cases.\n\nI saw great improvements of my LB scores since I started voting between submissions.\n\nWhen using voting strategies we also need to consider that voting can mask the underlying variability of each model’s performance on different data samples.\n\nIn other words, the ensemble might look consistent on the public set, but individual models may vary widely on certain data patterns, indicating that the model ensemble’s true generalization ability may be weaker than it appears.\n\nThis can have an impact in our case : Kaggle test splits (public vs. private) are typically designed to represent the distribution of the entire test set. However, there can still be subtle distributional shifts between the public and private splits due to random sampling.\n\nIf the ensemble voting strategy unintentionally overfits to specific quirks in the public test data, it may generalize poorly on the private test set. The final score on the private leaderboard could then be lower than expected.\n\nTo briefly conclude on this, voting currently leads to better leaderboard scores, but it doesn’t mean those scores will translate onto the final test set and leaderboard.",
    "3037309": "VotingRegressor doesn't mask anything - it returns the average of the results, it's just a fancy np.mean(your_results).\nHere is the official description - \"A voting regressor is an ensemble meta-estimator that fits several base regressors, each on the whole dataset. Then it averages the individual predictions to form a final prediction.\"",
    "3037321": "You’re absolutely right that VotingRegressor performs a straightforward averaging of predictions and doesn’t inherently mask variability. \n\nHowever, in this case, we’re not talking about VotingRegressors. Voting between submissions uses majority voting, which is different from averaging.\n\nBy using majority voting rather than averaging, the aim is to reduce the impact of outlier predictions from individual models. This approach is particularly helpful for discrete ordinal predictions, where taking the mode can reduce the influence of individual model biases, as it discards predictions that diverge too much from the majority. This is less likely to happen with simple averaging. \n\nI wish you good luck in this competition.",
    "3037335": "I think not always the autocoder can be the winner model.They might be the \"winner\" for tasks that rely on dimensionality reduction, feature extraction, or anomaly detection, especially when data is unlabeled.",
    "3037591": "I frequently observed the `GaussianNoise` layer receiving input data at the start of the AutoEncoder architectures, but here it's skipped, perhaps for some reason. On the step of scaling, I also noticed that the Gauss Rank Scaler was used rather than other scalers. Another addition is a Swap Noise. These three things I associate with Denoising AutoEncoder primarily. Here, as I can follow, AutoEncoder does feature extraction, so noise is not necessary, although it may be useful as a form of regularization to learn a more robust representation and improve generalization since there is a lack of data, but the issue of unidentified randomness still remains, thus how to test anything is a question.",
    "3037690": "sumit08 yes agree bout some portion of the data are unabled , so autoencoder may have the upper hand",
    "3041461": "Can you explain what is the mistake? thank you."
  },
  "source": "meta"
}