{
  "id": 548807,
  "title": "About using AutoEncoder",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/548807",
  "author_name": "",
  "post_date": "2024-11-29T01:33:32.955824Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've been looking at some notebooks with high scores, and a couple of things caught my eye. <br>\nThey're updating the AutoEncoder weights using both the training and test data, and they're also applying fit_transform() to the test data when standardizing.<br>\nIs this something people on Kaggle do intentionally for some reason? Or are they just copying code without fully understanding it?</p>",
  "messages": [
    {
      "id": "3058055",
      "postDate": "11/29/2024 01:33:32",
      "content": "<p>I've been looking at some notebooks with high scores, and a couple of things caught my eye. <br>\nThey're updating the AutoEncoder weights using both the training and test data, and they're also applying fit_transform() to the test data when standardizing.<br>\nIs this something people on Kaggle do intentionally for some reason? Or are they just copying code without fully understanding it?</p>",
      "rawMarkdown": "I've been looking at some notebooks with high scores, and a couple of things caught my eye. \nThey're updating the AutoEncoder weights using both the training and test data, and they're also applying fit_transform() to the test data when standardizing.\nIs this something people on Kaggle do intentionally for some reason? Or are they just copying code without fully understanding it?",
      "votes": null
    },
    {
      "id": "3058540",
      "postDate": "11/29/2024 15:09:34",
      "content": "<p>You are right. The usage of autoencoder in public notebooks are wrong and it causes severe data leakage, which may lead to an overfitting score on LB. Without enough training data, the performance of autoencoder is unstable. Up to now it seems that no one can get a high score on LB by correctly handling the missing values.</p>",
      "rawMarkdown": "You are right. The usage of autoencoder in public notebooks are wrong and it causes severe data leakage, which may lead to an overfitting score on LB. Without enough training data, the performance of autoencoder is unstable. Up to now it seems that no one can get a high score on LB by correctly handling the missing values.",
      "votes": null
    },
    {
      "id": "3059587",
      "postDate": "11/30/2024 21:28:40",
      "content": "<p>When I try just with the train set only the result is a little worse. Though this goes against conventional wisdom,<br>\nit might actually be capturing important distribution differences between train and test sets or maybe it is a form of data augmentation.<br>\nIt might be learning more robust or generalizable features?</p>",
      "rawMarkdown": "When I try just with the train set only the result is a little worse. Though this goes against conventional wisdom,\nit might actually be capturing important distribution differences between train and test sets or maybe it is a form of data augmentation.\nIt might be learning more robust or generalizable features?",
      "votes": null
    },
    {
      "id": "3062093",
      "postDate": "12/03/2024 09:32:01",
      "content": "<p>There are 2 types of users who uses it in this format:</p>\n<ol>\n<li>As simple as you said - \"they just copying code without fully understanding it\"</li>\n<li>Because the LB score is better like this - it's called overfitting the LB</li>\n</ol>\n<p>Both cases are not good practices, so use it correctly and wait for the final shake-up </p>",
      "rawMarkdown": "There are 2 types of users who uses it in this format:\n1. As simple as you said - \"they just copying code without fully understanding it\"\n2. Because the LB score is better like this - it's called overfitting the LB\n\nBoth cases are not good practices, so use it correctly and wait for the final shake-up",
      "votes": null
    },
    {
      "id": "3062842",
      "postDate": "12/04/2024 01:14:43",
      "content": "<p>Thanks for the reply. I now understand the reason. I hope eventually a good solution will be published.</p>",
      "rawMarkdown": "Thanks for the reply. I now understand the reason. I hope eventually a good solution will be published.",
      "votes": null
    },
    {
      "id": "3062844",
      "postDate": "12/04/2024 01:19:58",
      "content": "<p>Thanks for the reply. I hope the correct use of autoencoder will spread.</p>",
      "rawMarkdown": "Thanks for the reply. I hope the correct use of autoencoder will spread.",
      "votes": null
    },
    {
      "id": "3062852",
      "postDate": "12/04/2024 01:30:10",
      "content": "<p>Test data should be “unknown data”. If we update the weights with test data, the evaluation for that test data will no longer be fair.<br>\nHopefully a model will be developed that can be properly used for the healthy digital habits that this competition is aiming for.</p>",
      "rawMarkdown": "Test data should be “unknown data”. If we update the weights with test data, the evaluation for that test data will no longer be fair.\nHopefully a model will be developed that can be properly used for the healthy digital habits that this competition is aiming for.",
      "votes": null
    },
    {
      "id": "3075373",
      "postDate": "12/18/2024 18:52:00",
      "content": "<p>In fact, at that time there were several notebooks with the top scores and most of them used autoencoder, so I am pretty sure that it spread out quickly because of the high public score of those notebooks.</p>",
      "rawMarkdown": "In fact, at that time there were several notebooks with the top scores and most of them used autoencoder, so I am pretty sure that it spread out quickly because of the high public score of those notebooks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3058540,
      "author_name": "qufangcq",
      "author_url": "",
      "post_date": "11/29/2024 15:09:34",
      "content": "<p>You are right. The usage of autoencoder in public notebooks are wrong and it causes severe data leakage, which may lead to an overfitting score on LB. Without enough training data, the performance of autoencoder is unstable. Up to now it seems that no one can get a high score on LB by correctly handling the missing values.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3062844,
          "author_name": "kitazawakeiichi",
          "author_url": "",
          "post_date": "12/04/2024 01:19:58",
          "content": "<p>Thanks for the reply. I hope the correct use of autoencoder will spread.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3059587,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "11/30/2024 21:28:40",
      "content": "<p>When I try just with the train set only the result is a little worse. Though this goes against conventional wisdom,<br>\nit might actually be capturing important distribution differences between train and test sets or maybe it is a form of data augmentation.<br>\nIt might be learning more robust or generalizable features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3062852,
          "author_name": "kitazawakeiichi",
          "author_url": "",
          "post_date": "12/04/2024 01:30:10",
          "content": "<p>Test data should be “unknown data”. If we update the weights with test data, the evaluation for that test data will no longer be fair.<br>\nHopefully a model will be developed that can be properly used for the healthy digital habits that this competition is aiming for.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3062093,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "12/03/2024 09:32:01",
      "content": "<p>There are 2 types of users who uses it in this format:</p>\n<ol>\n<li>As simple as you said - \"they just copying code without fully understanding it\"</li>\n<li>Because the LB score is better like this - it's called overfitting the LB</li>\n</ol>\n<p>Both cases are not good practices, so use it correctly and wait for the final shake-up </p>",
      "votes": null,
      "replies": [
        {
          "id": 3062842,
          "author_name": "kitazawakeiichi",
          "author_url": "",
          "post_date": "12/04/2024 01:14:43",
          "content": "<p>Thanks for the reply. I now understand the reason. I hope eventually a good solution will be published.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3075373,
      "author_name": "viktoriamelkumyan",
      "author_url": "",
      "post_date": "12/18/2024 18:52:00",
      "content": "<p>In fact, at that time there were several notebooks with the top scores and most of them used autoencoder, so I am pretty sure that it spread out quickly because of the high public score of those notebooks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3058055": "I've been looking at some notebooks with high scores, and a couple of things caught my eye. \nThey're updating the AutoEncoder weights using both the training and test data, and they're also applying fit_transform() to the test data when standardizing.\nIs this something people on Kaggle do intentionally for some reason? Or are they just copying code without fully understanding it?",
    "3058540": "You are right. The usage of autoencoder in public notebooks are wrong and it causes severe data leakage, which may lead to an overfitting score on LB. Without enough training data, the performance of autoencoder is unstable. Up to now it seems that no one can get a high score on LB by correctly handling the missing values.",
    "3059587": "When I try just with the train set only the result is a little worse. Though this goes against conventional wisdom,\nit might actually be capturing important distribution differences between train and test sets or maybe it is a form of data augmentation.\nIt might be learning more robust or generalizable features?",
    "3062093": "There are 2 types of users who uses it in this format:\n1. As simple as you said - \"they just copying code without fully understanding it\"\n2. Because the LB score is better like this - it's called overfitting the LB\n\nBoth cases are not good practices, so use it correctly and wait for the final shake-up",
    "3062842": "Thanks for the reply. I now understand the reason. I hope eventually a good solution will be published.",
    "3062844": "Thanks for the reply. I hope the correct use of autoencoder will spread.",
    "3062852": "Test data should be “unknown data”. If we update the weights with test data, the evaluation for that test data will no longer be fair.\nHopefully a model will be developed that can be properly used for the healthy digital habits that this competition is aiming for.",
    "3075373": "In fact, at that time there were several notebooks with the top scores and most of them used autoencoder, so I am pretty sure that it spread out quickly because of the high public score of those notebooks."
  },
  "source": "meta"
}