{
  "id": 584796,
  "title": "20 predictor features found with all infinite values - either +ve or -ve",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584796",
  "author_name": "",
  "post_date": "2025-06-16T06:01:25.487145200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>There are around 20 predictor features found with  all infinite values: ['X697', 'X698', 'X699'] etc. in both train and test datasets. I need to understand if we have some reference average market values for same you can help with or shall we just drop these columns? Well dropping is a solution but then some of these features might be crucial for target label, so might lead to loss of information.</p>",
  "messages": [
    {
      "id": "3225222",
      "postDate": "06/16/2025 06:01:25",
      "content": "<p>There are around 20 predictor features found with  all infinite values: ['X697', 'X698', 'X699'] etc. in both train and test datasets. I need to understand if we have some reference average market values for same you can help with or shall we just drop these columns? Well dropping is a solution but then some of these features might be crucial for target label, so might lead to loss of information.</p>",
      "rawMarkdown": "There are around 20 predictor features found with  all infinite values: ['X697', 'X698', 'X699'] etc. in both train and test datasets. I need to understand if we have some reference average market values for same you can help with or shall we just drop these columns? Well dropping is a solution but then some of these features might be crucial for target label, so might lead to loss of information.",
      "votes": null
    },
    {
      "id": "3225247",
      "postDate": "06/16/2025 06:47:10",
      "content": "<p>In my opinion, simply ignoring these features is enough, since all of them in the test dataset are all '-inf' yet.</p>",
      "rawMarkdown": "In my opinion, simply ignoring these features is enough, since all of them in the test dataset are all '-inf' yet.",
      "votes": null
    },
    {
      "id": "3225299",
      "postDate": "06/16/2025 08:27:46",
      "content": "<p>Sure, also my session is crashing even though I have imported only most recent less than half of data (250,000 records). Just want to know how many datapoints you have considered in your problem without data loss.</p>",
      "rawMarkdown": "Sure, also my session is crashing even though I have imported only most recent less than half of data (250,000 records). Just want to know how many datapoints you have considered in your problem without data loss.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3225247,
      "author_name": "littlecitizen",
      "author_url": "",
      "post_date": "06/16/2025 06:47:10",
      "content": "<p>In my opinion, simply ignoring these features is enough, since all of them in the test dataset are all '-inf' yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3225299,
      "author_name": "kumanchi",
      "author_url": "",
      "post_date": "06/16/2025 08:27:46",
      "content": "<p>Sure, also my session is crashing even though I have imported only most recent less than half of data (250,000 records). Just want to know how many datapoints you have considered in your problem without data loss.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3225222": "There are around 20 predictor features found with  all infinite values: ['X697', 'X698', 'X699'] etc. in both train and test datasets. I need to understand if we have some reference average market values for same you can help with or shall we just drop these columns? Well dropping is a solution but then some of these features might be crucial for target label, so might lead to loss of information.",
    "3225247": "In my opinion, simply ignoring these features is enough, since all of them in the test dataset are all '-inf' yet.",
    "3225299": "Sure, also my session is crashing even though I have imported only most recent less than half of data (250,000 records). Just want to know how many datapoints you have considered in your problem without data loss."
  },
  "source": "meta"
}