{
  "id": 540078,
  "title": "Combined Dataset for CSV and Parquet files",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/540078",
  "author_name": "",
  "post_date": "2024-10-12T13:16:37.530213400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>For testing different techniques, I have to run the notebook multiple times and every time I run the notebook I had to wait for the parquet files to be processed and loaded. To save time I combined the data and saved them as CSV.<br>\nSo, I have combined both csv and parquet for convenience. <br>\n<strong>Note:</strong><br>\nUse combined version only for testing purposes. Don't use it for submission to competition because after submission our code is run with unseen data in this competition.<br>\nTo use and view the combined data set for testing purposes you can visit <a href=\"https://www.kaggle.com/datasets/taimour/cmi-problematic-internet-usage-combined-data/data\" target=\"_blank\">Dataset here</a> </p>",
  "messages": [
    {
      "id": "3015475",
      "postDate": "10/12/2024 13:16:37",
      "content": "<p>For testing different techniques, I have to run the notebook multiple times and every time I run the notebook I had to wait for the parquet files to be processed and loaded. To save time I combined the data and saved them as CSV.<br>\nSo, I have combined both csv and parquet for convenience. <br>\n<strong>Note:</strong><br>\nUse combined version only for testing purposes. Don't use it for submission to competition because after submission our code is run with unseen data in this competition.<br>\nTo use and view the combined data set for testing purposes you can visit <a href=\"https://www.kaggle.com/datasets/taimour/cmi-problematic-internet-usage-combined-data/data\" target=\"_blank\">Dataset here</a> </p>",
      "rawMarkdown": "For testing different techniques, I have to run the notebook multiple times and every time I run the notebook I had to wait for the parquet files to be processed and loaded. To save time I combined the data and saved them as CSV.\nSo, I have combined both csv and parquet for convenience. \n\n**Note:**\nUse combined version only for testing purposes. Don't use it for submission to competition because after submission our code is run with unseen data in this competition.\n\n\nTo use and view the combined data set for testing purposes you can visit [Dataset here](https://www.kaggle.com/datasets/taimour/cmi-problematic-internet-usage-combined-data/data)",
      "votes": null
    },
    {
      "id": "3015605",
      "postDate": "10/12/2024 16:51:14",
      "content": "<p>are you using a special accelerator or GPU to load the parquet files? I am running into memory limitations loading the parquet files in the notebook</p>",
      "rawMarkdown": "are you using a special accelerator or GPU to load the parquet files? I am running into memory limitations loading the parquet files in the notebook",
      "votes": null
    },
    {
      "id": "3015611",
      "postDate": "10/12/2024 17:00:16",
      "content": "<p>I have loaded parquet files without an accelerator and I have also loaded parquet files with an accelerator both works fine. There is no problem of memory limitations.</p>\n<p>My code for loading parquet files with out GPU = <a href=\"https://www.kaggle.com/code/taimour/xgb-gridsearchcv-problematic-internet-usage\" target=\"_blank\">🧑‍💻 XGB GridSearchCV Problematic Internet Usage</a> </p>\n<p>My code for loading parquet files with GPU = <a href=\"https://www.kaggle.com/code/taimour/blend-xgb-lgbm-eda-cmi-internet-usage\" target=\"_blank\">💻 Blend XGB LGBM - EDA 📊 - CMI Internet Usage</a></p>",
      "rawMarkdown": "I have loaded parquet files without an accelerator and I have also loaded parquet files with an accelerator both works fine. There is no problem of memory limitations.\n\nMy code for loading parquet files with out GPU = [🧑‍💻 XGB GridSearchCV Problematic Internet Usage](https://www.kaggle.com/code/taimour/xgb-gridsearchcv-problematic-internet-usage) \n\nMy code for loading parquet files with GPU = [💻 Blend XGB LGBM - EDA 📊 - CMI Internet Usage](https://www.kaggle.com/code/taimour/blend-xgb-lgbm-eda-cmi-internet-usage)",
      "votes": null
    },
    {
      "id": "3016258",
      "postDate": "10/13/2024 14:50:07",
      "content": "<p>Did you apply dimensionality reduction before training?</p>",
      "rawMarkdown": "Did you apply dimensionality reduction before training?",
      "votes": null
    },
    {
      "id": "3017152",
      "postDate": "10/14/2024 15:07:12",
      "content": "<p>This dataset only combines csv and parquet files. I didn't applied dimensionality reduction while preparing these csv files.</p>",
      "rawMarkdown": "This dataset only combines csv and parquet files. I didn't applied dimensionality reduction while preparing these csv files.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3015605,
      "author_name": "f0ster",
      "author_url": "",
      "post_date": "10/12/2024 16:51:14",
      "content": "<p>are you using a special accelerator or GPU to load the parquet files? I am running into memory limitations loading the parquet files in the notebook</p>",
      "votes": null,
      "replies": [
        {
          "id": 3015611,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "10/12/2024 17:00:16",
          "content": "<p>I have loaded parquet files without an accelerator and I have also loaded parquet files with an accelerator both works fine. There is no problem of memory limitations.</p>\n<p>My code for loading parquet files with out GPU = <a href=\"https://www.kaggle.com/code/taimour/xgb-gridsearchcv-problematic-internet-usage\" target=\"_blank\">🧑‍💻 XGB GridSearchCV Problematic Internet Usage</a> </p>\n<p>My code for loading parquet files with GPU = <a href=\"https://www.kaggle.com/code/taimour/blend-xgb-lgbm-eda-cmi-internet-usage\" target=\"_blank\">💻 Blend XGB LGBM - EDA 📊 - CMI Internet Usage</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3016258,
      "author_name": "manojajj",
      "author_url": "",
      "post_date": "10/13/2024 14:50:07",
      "content": "<p>Did you apply dimensionality reduction before training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3017152,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "10/14/2024 15:07:12",
          "content": "<p>This dataset only combines csv and parquet files. I didn't applied dimensionality reduction while preparing these csv files.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3015475": "For testing different techniques, I have to run the notebook multiple times and every time I run the notebook I had to wait for the parquet files to be processed and loaded. To save time I combined the data and saved them as CSV.\nSo, I have combined both csv and parquet for convenience. \n\n**Note:**\nUse combined version only for testing purposes. Don't use it for submission to competition because after submission our code is run with unseen data in this competition.\n\n\nTo use and view the combined data set for testing purposes you can visit [Dataset here](https://www.kaggle.com/datasets/taimour/cmi-problematic-internet-usage-combined-data/data)",
    "3015605": "are you using a special accelerator or GPU to load the parquet files? I am running into memory limitations loading the parquet files in the notebook",
    "3015611": "I have loaded parquet files without an accelerator and I have also loaded parquet files with an accelerator both works fine. There is no problem of memory limitations.\n\nMy code for loading parquet files with out GPU = [🧑‍💻 XGB GridSearchCV Problematic Internet Usage](https://www.kaggle.com/code/taimour/xgb-gridsearchcv-problematic-internet-usage) \n\nMy code for loading parquet files with GPU = [💻 Blend XGB LGBM - EDA 📊 - CMI Internet Usage](https://www.kaggle.com/code/taimour/blend-xgb-lgbm-eda-cmi-internet-usage)",
    "3016258": "Did you apply dimensionality reduction before training?",
    "3017152": "This dataset only combines csv and parquet files. I didn't applied dimensionality reduction while preparing these csv files."
  },
  "source": "meta"
}