{
  "id": 548340,
  "title": "How to handle the parquet information? Does it make sense to train a seperate model on it?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/548340",
  "author_name": "",
  "post_date": "2024-11-26T08:38:32.396740600Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nI wondered if it would make sense to train a seperate model just on the parquet information and/or without its matching csv part. My reasoning is that by training a model on the full df (also csv without parquet information), the parquet information will never have a high feature importance no matter how good you feature engineer it since it either 2/3rd of your df misses this data or (if you use imputation) you probably erase any signal it has.<br>\nSo my ideas where either to train a seperate model just on the ~900 data points and then overwrite the sii from the model with the complete df. Or train a model just on the parquet part and sii information and create a new feature with the result for the model with the complete df.<br>\nDoes this make sense? How do you normally handle this?<br>\nI assume it will not help alot, since it is probably hard to extract a good signal from those files. But I'm still interested what would be your normal approach if the information wasn't that noisy.</p>",
  "messages": [
    {
      "id": "3055882",
      "postDate": "11/26/2024 08:38:32",
      "content": "<p>Hello everyone,<br>\nI wondered if it would make sense to train a seperate model just on the parquet information and/or without its matching csv part. My reasoning is that by training a model on the full df (also csv without parquet information), the parquet information will never have a high feature importance no matter how good you feature engineer it since it either 2/3rd of your df misses this data or (if you use imputation) you probably erase any signal it has.<br>\nSo my ideas where either to train a seperate model just on the ~900 data points and then overwrite the sii from the model with the complete df. Or train a model just on the parquet part and sii information and create a new feature with the result for the model with the complete df.<br>\nDoes this make sense? How do you normally handle this?<br>\nI assume it will not help alot, since it is probably hard to extract a good signal from those files. But I'm still interested what would be your normal approach if the information wasn't that noisy.</p>",
      "rawMarkdown": "Hello everyone,\nI wondered if it would make sense to train a seperate model just on the parquet information and/or without its matching csv part. My reasoning is that by training a model on the full df (also csv without parquet information), the parquet information will never have a high feature importance no matter how good you feature engineer it since it either 2/3rd of your df misses this data or (if you use imputation) you probably erase any signal it has.\nSo my ideas where either to train a seperate model just on the ~900 data points and then overwrite the sii from the model with the complete df. Or train a model just on the parquet part and sii information and create a new feature with the result for the model with the complete df.\nDoes this make sense? How do you normally handle this?\nI assume it will not help alot, since it is probably hard to extract a good signal from those files. But I'm still interested what would be your normal approach if the information wasn't that noisy.",
      "votes": null
    },
    {
      "id": "3063620",
      "postDate": "12/04/2024 17:06:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> , I'd heard of this idea in another discussion and it seemed/seems like a good idea. Also, I've reduced the 996 ids with parquet data down to around 800 that have at least a week of good data, so that's more reason to have a specialized model.</p>\n<p>So far I haven't seen an improvement with the separate model, but I think that's mostly because my parquet features are not very useful in predicting the Total/sii values. Wishing you better luck/skill 🙂</p>",
      "rawMarkdown": "Hi @mariusheuser , I'd heard of this idea in another discussion and it seemed/seems like a good idea. Also, I've reduced the 996 ids with parquet data down to around 800 that have at least a week of good data, so that's more reason to have a specialized model.\n\nSo far I haven't seen an improvement with the separate model, but I think that's mostly because my parquet features are not very useful in predicting the Total/sii values. Wishing you better luck/skill 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3063620,
      "author_name": "dan3dewey",
      "author_url": "",
      "post_date": "12/04/2024 17:06:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> , I'd heard of this idea in another discussion and it seemed/seems like a good idea. Also, I've reduced the 996 ids with parquet data down to around 800 that have at least a week of good data, so that's more reason to have a specialized model.</p>\n<p>So far I haven't seen an improvement with the separate model, but I think that's mostly because my parquet features are not very useful in predicting the Total/sii values. Wishing you better luck/skill 🙂</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3055882": "Hello everyone,\nI wondered if it would make sense to train a seperate model just on the parquet information and/or without its matching csv part. My reasoning is that by training a model on the full df (also csv without parquet information), the parquet information will never have a high feature importance no matter how good you feature engineer it since it either 2/3rd of your df misses this data or (if you use imputation) you probably erase any signal it has.\nSo my ideas where either to train a seperate model just on the ~900 data points and then overwrite the sii from the model with the complete df. Or train a model just on the parquet part and sii information and create a new feature with the result for the model with the complete df.\nDoes this make sense? How do you normally handle this?\nI assume it will not help alot, since it is probably hard to extract a good signal from those files. But I'm still interested what would be your normal approach if the information wasn't that noisy.",
    "3063620": "Hi @mariusheuser , I'd heard of this idea in another discussion and it seemed/seems like a good idea. Also, I've reduced the 996 ids with parquet data down to around 800 that have at least a week of good data, so that's more reason to have a specialized model.\n\nSo far I haven't seen an improvement with the separate model, but I think that's mostly because my parquet features are not very useful in predicting the Total/sii values. Wishing you better luck/skill 🙂"
  },
  "source": "meta"
}