{
  "id": 580468,
  "title": "How to Approach EDA with a lot of Anonymized Features?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580468",
  "author_name": "",
  "post_date": "2025-05-24T07:07:15.173201900Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>As we have noticed that the dataset contains more then 800 features and almost all the features don't have meaningful names. I am wondering how others are approaching EDA in this case. <br>\n<strong>What strategies are you using to explore and make sense of such a large set of unnamed features? Any tips on identifying useful signals or reducing dimensionality?</strong></p>\n<p>Looking forward to suggestions from experts and anyone who's dealt with similar scenarios before.</p>",
  "messages": [
    {
      "id": "3208449",
      "postDate": "05/24/2025 07:07:15",
      "content": "<p>Hi Everyone,</p>\n<p>As we have noticed that the dataset contains more then 800 features and almost all the features don't have meaningful names. I am wondering how others are approaching EDA in this case. <br>\n<strong>What strategies are you using to explore and make sense of such a large set of unnamed features? Any tips on identifying useful signals or reducing dimensionality?</strong></p>\n<p>Looking forward to suggestions from experts and anyone who's dealt with similar scenarios before.</p>",
      "rawMarkdown": "Hi Everyone,\n\nAs we have noticed that the dataset contains more then 800 features and almost all the features don't have meaningful names. I am wondering how others are approaching EDA in this case. \n**What strategies are you using to explore and make sense of such a large set of unnamed features? Any tips on identifying useful signals or reducing dimensionality?**\n\nLooking forward to suggestions from experts and anyone who's dealt with similar scenarios before.",
      "votes": null
    },
    {
      "id": "3210094",
      "postDate": "05/26/2025 18:01:15",
      "content": "<p>This might be helpful. Try the following methods.</p>\n<ul>\n<li>Under the Nature of the Features</li>\n<li>Dimensionality Reduction</li>\n<li>Check feature Importance</li>\n<li>Can use feature Clustering</li>\n</ul>",
      "rawMarkdown": "This might be helpful. Try the following methods.\n- Under the Nature of the Features\n- Dimensionality Reduction\n- Check feature Importance\n- Can use feature Clustering",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3210094,
      "author_name": "sharmajicoder",
      "author_url": "",
      "post_date": "05/26/2025 18:01:15",
      "content": "<p>This might be helpful. Try the following methods.</p>\n<ul>\n<li>Under the Nature of the Features</li>\n<li>Dimensionality Reduction</li>\n<li>Check feature Importance</li>\n<li>Can use feature Clustering</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3208449": "Hi Everyone,\n\nAs we have noticed that the dataset contains more then 800 features and almost all the features don't have meaningful names. I am wondering how others are approaching EDA in this case. \n**What strategies are you using to explore and make sense of such a large set of unnamed features? Any tips on identifying useful signals or reducing dimensionality?**\n\nLooking forward to suggestions from experts and anyone who's dealt with similar scenarios before.",
    "3210094": "This might be helpful. Try the following methods.\n- Under the Nature of the Features\n- Dimensionality Reduction\n- Check feature Importance\n- Can use feature Clustering"
  },
  "source": "meta"
}