{
  "id": 546499,
  "title": "Single stat or Multiple stats? What is best for parquet files?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546499",
  "author_name": "Taimour Nazar",
  "post_date": "2024-11-16T08:15:07.955000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello</p>\n<p>Looking at the current state of data we know that it is noisy, less in quantity and has missing values. So, should we compute descriptive statistics using <code>df.describe().values</code> or should we use only <code>df.mean().values</code>  which will be a single stat? (mean or any other single stat)</p>\n<p>Majority of people or we can say almost everyone is using <code>df.describe().values</code> don't you think that this way of handling the data introduces more noise in the data? Will using any single state prove to be better in our scenario?</p>\n<p>Although I have noticed any big difference so far in my notebook <a href=\"https://www.kaggle.com/code/taimour/blend-xgb-lgbm-cat-in-depth-eda-cmi\" target=\"_blank\">💻 Blend XGB LGBM CAT - 📊In Depth EDA | CMI</a> i.e in version 31, 32, 33</p>",
  "messages": [
    {
      "id": 3047077,
      "postDate": "2024-11-16T08:15:07.957Z",
      "content": "<p>Hello</p>\n<p>Looking at the current state of data we know that it is noisy, less in quantity and has missing values. So, should we compute descriptive statistics using <code>df.describe().values</code> or should we use only <code>df.mean().values</code>  which will be a single stat? (mean or any other single stat)</p>\n<p>Majority of people or we can say almost everyone is using <code>df.describe().values</code> don't you think that this way of handling the data introduces more noise in the data? Will using any single state prove to be better in our scenario?</p>\n<p>Although I have noticed any big difference so far in my notebook <a href=\"https://www.kaggle.com/code/taimour/blend-xgb-lgbm-cat-in-depth-eda-cmi\" target=\"_blank\">💻 Blend XGB LGBM CAT - 📊In Depth EDA | CMI</a> i.e in version 31, 32, 33</p>",
      "rawMarkdown": "Hello\n\nLooking at the current state of data we know that it is noisy, less in quantity and has missing values. So, should we compute descriptive statistics using `df.describe().values` or should we use only `df.mean().values`  which will be a single stat? (mean or any other single stat)\n\nMajority of people or we can say almost everyone is using `df.describe().values` don't you think that this way of handling the data introduces more noise in the data? Will using any single state prove to be better in our scenario?\n\nAlthough I have noticed any big difference so far in my notebook [💻 Blend XGB LGBM CAT - 📊In Depth EDA | CMI](https://www.kaggle.com/code/taimour/blend-xgb-lgbm-cat-in-depth-eda-cmi) i.e in version 31, 32, 33",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3047077": "Hello\n\nLooking at the current state of data we know that it is noisy, less in quantity and has missing values. So, should we compute descriptive statistics using `df.describe().values` or should we use only `df.mean().values`  which will be a single stat? (mean or any other single stat)\n\nMajority of people or we can say almost everyone is using `df.describe().values` don't you think that this way of handling the data introduces more noise in the data? Will using any single state prove to be better in our scenario?\n\nAlthough I have noticed any big difference so far in my notebook [💻 Blend XGB LGBM CAT - 📊In Depth EDA | CMI](https://www.kaggle.com/code/taimour/blend-xgb-lgbm-cat-in-depth-eda-cmi) i.e in version 31, 32, 33"
  }
}