{
  "id": 551202,
  "title": "Any useful features from the parquet time series?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551202",
  "author_name": "",
  "post_date": "2024-12-11T21:48:48.123340300Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm curious if anyone created some useful features using the parquet time series data, beyond just using the descriptive statistics?</p>\n<p>I ask because I created some that were marginally useful, for examples: i) average bedtime, ii) average hours of sleep, iii) average enmo when awake, and iv) the overlap (average of the products) of enmo and light on the weekends. These are significantly correlated with PCIAT Total and have some Permutation Importance to the XGB model. Adding the latter feature (iv) to a simple model gave an LB increase of maybe 0.010.</p>\n<p>Because there are useful parquet files for only about 1/3 of the ids, I also tried doing a second model just on them and then patched its predictions into the All-ids model predictions. This was promising on the GSCV \"test\" scores, but submissions always had lower LB scores by about 0.015.</p>\n<p>Let me/us know if you had more success - without giving away your \"secret sauce\" recipe, of course 😉</p>\n<p>fyi, Here are links to a couple output images from my <a href=\"https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction\" target=\"_blank\">CMI 2024 - Kaggle Addiction ?</a> notebook: <a href=\"https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Sleep_Awake_times.png\" target=\"_blank\">Sleep and Wake Times</a> and <a href=\"https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Total_vs_wkend_enmoXlight.png\" target=\"_blank\">Total vs enmoXlight feature</a></p>",
  "messages": [
    {
      "id": "3069781",
      "postDate": "12/11/2024 21:48:48",
      "content": "<p>I'm curious if anyone created some useful features using the parquet time series data, beyond just using the descriptive statistics?</p>\n<p>I ask because I created some that were marginally useful, for examples: i) average bedtime, ii) average hours of sleep, iii) average enmo when awake, and iv) the overlap (average of the products) of enmo and light on the weekends. These are significantly correlated with PCIAT Total and have some Permutation Importance to the XGB model. Adding the latter feature (iv) to a simple model gave an LB increase of maybe 0.010.</p>\n<p>Because there are useful parquet files for only about 1/3 of the ids, I also tried doing a second model just on them and then patched its predictions into the All-ids model predictions. This was promising on the GSCV \"test\" scores, but submissions always had lower LB scores by about 0.015.</p>\n<p>Let me/us know if you had more success - without giving away your \"secret sauce\" recipe, of course 😉</p>\n<p>fyi, Here are links to a couple output images from my <a href=\"https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction\" target=\"_blank\">CMI 2024 - Kaggle Addiction ?</a> notebook: <a href=\"https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Sleep_Awake_times.png\" target=\"_blank\">Sleep and Wake Times</a> and <a href=\"https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Total_vs_wkend_enmoXlight.png\" target=\"_blank\">Total vs enmoXlight feature</a></p>",
      "rawMarkdown": "I'm curious if anyone created some useful features using the parquet time series data, beyond just using the descriptive statistics?\n\nI ask because I created some that were marginally useful, for examples: i) average bedtime, ii) average hours of sleep, iii) average enmo when awake, and iv) the overlap (average of the products) of enmo and light on the weekends. These are significantly correlated with PCIAT Total and have some Permutation Importance to the XGB model. Adding the latter feature (iv) to a simple model gave an LB increase of maybe 0.010.\n\nBecause there are useful parquet files for only about 1/3 of the ids, I also tried doing a second model just on them and then patched its predictions into the All-ids model predictions. This was promising on the GSCV \"test\" scores, but submissions always had lower LB scores by about 0.015.\n\nLet me/us know if you had more success - without giving away your \"secret sauce\" recipe, of course 😉\n\nfyi, Here are links to a couple output images from my [CMI 2024 - Kaggle Addiction ?](https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction) notebook: [Sleep and Wake Times](https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Sleep_Awake_times.png) and [Total vs enmoXlight feature](https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Total_vs_wkend_enmoXlight.png)",
      "votes": null
    },
    {
      "id": "3069807",
      "postDate": "12/11/2024 23:01:04",
      "content": "<p>In descriptive statistics, stat_90 was the only interesting feature and could give some ideas.</p>\n<p>Thank you for having shared your notebook. I can not share mine yet. </p>\n<p>I tried more than 900 actipraphy features and I had success with less than 15 of them ; i'm using permutation importance too. And my actigraphy features improve my CV score from 0.02.</p>",
      "rawMarkdown": "In descriptive statistics, stat_90 was the only interesting feature and could give some ideas.\n\nThank you for having shared your notebook. I can not share mine yet. \n\nI tried more than 900 actipraphy features and I had success with less than 15 of them ; i'm using permutation importance too. And my actigraphy features improve my CV score from 0.02.",
      "votes": null
    },
    {
      "id": "3069814",
      "postDate": "12/11/2024 23:27:46",
      "content": "<p>Like you mentioned, I also tried alot of different features but my CV didn't change alot.<br>\nThe most meaningful for my models was strangely enough max_light, which you could argue is helpful. But I didnt really clean the parquet file and there are alot of unrealistic values, so max isn't the best measurement. But it still performs better than the others for my models…</p>\n<p>I have another question. I engineered alot of parquet features. What is the best approach now?<br>\nAutoencode them to reduce the amount of features or reduce them via feature importance?</p>",
      "rawMarkdown": "Like you mentioned, I also tried alot of different features but my CV didn't change alot.\nThe most meaningful for my models was strangely enough max_light, which you could argue is helpful. But I didnt really clean the parquet file and there are alot of unrealistic values, so max isn't the best measurement. But it still performs better than the others for my models...\n\nI have another question. I engineered alot of parquet features. What is the best approach now?\nAutoencode them to reduce the amount of features or reduce them via feature importance?",
      "votes": null
    },
    {
      "id": "3069943",
      "postDate": "12/12/2024 04:35:21",
      "content": "<p>Same here- most of the actigraphy features are totally useless. I am quite confused about choosing my final submission now <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> </p>",
      "rawMarkdown": "Same here- most of the actigraphy features are totally useless. I am quite confused about choosing my final submission now @adaubas",
      "votes": null
    },
    {
      "id": "3069952",
      "postDate": "12/12/2024 04:56:25",
      "content": "<blockquote>\n  <p>Autoencode them to reduce the amount of features or reduce them via feature importance?</p>\n</blockquote>\n<p>This is totally your call, but I can only suggest that you may refrain from reading and incorporating public ideas now at this stage of the competition and rather rely on your experiments instead <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> </p>",
      "rawMarkdown": "> Autoencode them to reduce the amount of features or reduce them via feature importance?\n\nThis is totally your call, but I can only suggest that you may refrain from reading and incorporating public ideas now at this stage of the competition and rather rely on your experiments instead @mariusheuser",
      "votes": null
    },
    {
      "id": "3069955",
      "postDate": "12/12/2024 04:58:13",
      "content": "<p>AutoEncoder is not trained with target. Feature Importance will select features which are important for the target.</p>",
      "rawMarkdown": "AutoEncoder is not trained with target. Feature Importance will select features which are important for the target.",
      "votes": null
    },
    {
      "id": "3070053",
      "postDate": "12/12/2024 07:17:21",
      "content": "<p>I made custom NN to encode time series data.<br>\nOnly NN single model gets almost same CV/LB score as gbdt model.<br>\nI will make it public after end of comp. </p>",
      "rawMarkdown": "I made custom NN to encode time series data.\nOnly NN single model gets almost same CV/LB score as gbdt model.\nI will make it public after end of comp.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3069807,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/11/2024 23:01:04",
      "content": "<p>In descriptive statistics, stat_90 was the only interesting feature and could give some ideas.</p>\n<p>Thank you for having shared your notebook. I can not share mine yet. </p>\n<p>I tried more than 900 actipraphy features and I had success with less than 15 of them ; i'm using permutation importance too. And my actigraphy features improve my CV score from 0.02.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3069943,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/12/2024 04:35:21",
          "content": "<p>Same here- most of the actigraphy features are totally useless. I am quite confused about choosing my final submission now <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3069814,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "12/11/2024 23:27:46",
      "content": "<p>Like you mentioned, I also tried alot of different features but my CV didn't change alot.<br>\nThe most meaningful for my models was strangely enough max_light, which you could argue is helpful. But I didnt really clean the parquet file and there are alot of unrealistic values, so max isn't the best measurement. But it still performs better than the others for my models…</p>\n<p>I have another question. I engineered alot of parquet features. What is the best approach now?<br>\nAutoencode them to reduce the amount of features or reduce them via feature importance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3069952,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/12/2024 04:56:25",
          "content": "<blockquote>\n  <p>Autoencode them to reduce the amount of features or reduce them via feature importance?</p>\n</blockquote>\n<p>This is totally your call, but I can only suggest that you may refrain from reading and incorporating public ideas now at this stage of the competition and rather rely on your experiments instead <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3069955,
          "author_name": "adaubas",
          "author_url": "",
          "post_date": "12/12/2024 04:58:13",
          "content": "<p>AutoEncoder is not trained with target. Feature Importance will select features which are important for the target.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3070053,
      "author_name": "clearwaterkzk",
      "author_url": "",
      "post_date": "12/12/2024 07:17:21",
      "content": "<p>I made custom NN to encode time series data.<br>\nOnly NN single model gets almost same CV/LB score as gbdt model.<br>\nI will make it public after end of comp. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3069781": "I'm curious if anyone created some useful features using the parquet time series data, beyond just using the descriptive statistics?\n\nI ask because I created some that were marginally useful, for examples: i) average bedtime, ii) average hours of sleep, iii) average enmo when awake, and iv) the overlap (average of the products) of enmo and light on the weekends. These are significantly correlated with PCIAT Total and have some Permutation Importance to the XGB model. Adding the latter feature (iv) to a simple model gave an LB increase of maybe 0.010.\n\nBecause there are useful parquet files for only about 1/3 of the ids, I also tried doing a second model just on them and then patched its predictions into the All-ids model predictions. This was promising on the GSCV \"test\" scores, but submissions always had lower LB scores by about 0.015.\n\nLet me/us know if you had more success - without giving away your \"secret sauce\" recipe, of course 😉\n\nfyi, Here are links to a couple output images from my [CMI 2024 - Kaggle Addiction ?](https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction) notebook: [Sleep and Wake Times](https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Sleep_Awake_times.png) and [Total vs enmoXlight feature](https://www.kaggle.com/code/dan3dewey/cmi-2024-kaggle-addiction/output?select=Total_vs_wkend_enmoXlight.png)",
    "3069807": "In descriptive statistics, stat_90 was the only interesting feature and could give some ideas.\n\nThank you for having shared your notebook. I can not share mine yet. \n\nI tried more than 900 actipraphy features and I had success with less than 15 of them ; i'm using permutation importance too. And my actigraphy features improve my CV score from 0.02.",
    "3069814": "Like you mentioned, I also tried alot of different features but my CV didn't change alot.\nThe most meaningful for my models was strangely enough max_light, which you could argue is helpful. But I didnt really clean the parquet file and there are alot of unrealistic values, so max isn't the best measurement. But it still performs better than the others for my models...\n\nI have another question. I engineered alot of parquet features. What is the best approach now?\nAutoencode them to reduce the amount of features or reduce them via feature importance?",
    "3069943": "Same here- most of the actigraphy features are totally useless. I am quite confused about choosing my final submission now @adaubas",
    "3069952": "> Autoencode them to reduce the amount of features or reduce them via feature importance?\n\nThis is totally your call, but I can only suggest that you may refrain from reading and incorporating public ideas now at this stage of the competition and rather rely on your experiments instead @mariusheuser",
    "3069955": "AutoEncoder is not trained with target. Feature Importance will select features which are important for the target.",
    "3070053": "I made custom NN to encode time series data.\nOnly NN single model gets almost same CV/LB score as gbdt model.\nI will make it public after end of comp."
  },
  "source": "meta"
}