{
  "id": 361656,
  "title": "Help in EDA and downloading the dataset",
  "url": "/competitions/open-problems-multimodal/discussion/361656",
  "author_name": "",
  "post_date": "2022-10-22T22:22:02.102475500Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey, I am a beginner in Computer Vision.  I don't know how to process this huge size of data and how to perform  EDA( I have done all this on a relatively very small dataset). Any help will be highly appreciated. Thanks </p>",
  "messages": [
    {
      "id": "1999963",
      "postDate": "10/22/2022 22:22:02",
      "content": "<p>Hey, I am a beginner in Computer Vision.  I don't know how to process this huge size of data and how to perform  EDA( I have done all this on a relatively very small dataset). Any help will be highly appreciated. Thanks </p>",
      "rawMarkdown": "Hey, I am a beginner in Computer Vision.  I don't know how to process this huge size of data and how to perform  EDA( I have done all this on a relatively very small dataset). Any help will be highly appreciated. Thanks",
      "votes": null
    },
    {
      "id": "2000191",
      "postDate": "10/23/2022 05:55:04",
      "content": "<p>The dataset present in this competition are in h5 format. For initial analysis you can just load part of dataset.<br>\nFor example, pd.read_hdf(data, start=0, stop=5000), this will load only the first 5000 lines.<br>\nFor working on the whole dataset, the best approach is to convert the dataset into sparse matrix format. WHY? because the CITE-seq data has around 75% zeros and ATAC-seq has around 98% zeros. This notebook : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense</a> , will help you on EDA front.</p>",
      "rawMarkdown": "The dataset present in this competition are in h5 format. For initial analysis you can just load part of dataset.\nFor example, pd.read_hdf(data, start=0, stop=5000), this will load only the first 5000 lines.\nFor working on the whole dataset, the best approach is to convert the dataset into sparse matrix format. WHY? because the CITE-seq data has around 75% zeros and ATAC-seq has around 98% zeros. This notebook : [https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense](url) , will help you on EDA front.",
      "votes": null
    },
    {
      "id": "2000201",
      "postDate": "10/23/2022 06:08:43",
      "content": "<p>Thanks for the help. Appreciate it</p>",
      "rawMarkdown": "Thanks for the help. Appreciate it",
      "votes": null
    },
    {
      "id": "2000684",
      "postDate": "10/23/2022 13:40:48",
      "content": "<p>As <a href=\"https://www.kaggle.com/chandanpandey\" target=\"_blank\">@chandanpandey</a> pointed out, sparse data structures are the way to work with this data.  I put together <a href=\"https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets/notebook\" target=\"_blank\">this notebook</a> to provide some code to read the data in chunks, identify columns that have low variance (there are some columns that are all zeros), and a class to capture all those details.</p>",
      "rawMarkdown": "As @chandanpandey pointed out, sparse data structures are the way to work with this data.  I put together [this notebook](https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets/notebook) to provide some code to read the data in chunks, identify columns that have low variance (there are some columns that are all zeros), and a class to capture all those details.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2000191,
      "author_name": "chandanpandey",
      "author_url": "",
      "post_date": "10/23/2022 05:55:04",
      "content": "<p>The dataset present in this competition are in h5 format. For initial analysis you can just load part of dataset.<br>\nFor example, pd.read_hdf(data, start=0, stop=5000), this will load only the first 5000 lines.<br>\nFor working on the whole dataset, the best approach is to convert the dataset into sparse matrix format. WHY? because the CITE-seq data has around 75% zeros and ATAC-seq has around 98% zeros. This notebook : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense</a> , will help you on EDA front.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2000201,
          "author_name": "takihasan",
          "author_url": "",
          "post_date": "10/23/2022 06:08:43",
          "content": "<p>Thanks for the help. Appreciate it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2000684,
      "author_name": "kirkdco",
      "author_url": "",
      "post_date": "10/23/2022 13:40:48",
      "content": "<p>As <a href=\"https://www.kaggle.com/chandanpandey\" target=\"_blank\">@chandanpandey</a> pointed out, sparse data structures are the way to work with this data.  I put together <a href=\"https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets/notebook\" target=\"_blank\">this notebook</a> to provide some code to read the data in chunks, identify columns that have low variance (there are some columns that are all zeros), and a class to capture all those details.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1999963": "Hey, I am a beginner in Computer Vision.  I don't know how to process this huge size of data and how to perform  EDA( I have done all this on a relatively very small dataset). Any help will be highly appreciated. Thanks",
    "2000191": "The dataset present in this competition are in h5 format. For initial analysis you can just load part of dataset.\nFor example, pd.read_hdf(data, start=0, stop=5000), this will load only the first 5000 lines.\nFor working on the whole dataset, the best approach is to convert the dataset into sparse matrix format. WHY? because the CITE-seq data has around 75% zeros and ATAC-seq has around 98% zeros. This notebook : [https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense](url) , will help you on EDA front.",
    "2000201": "Thanks for the help. Appreciate it",
    "2000684": "As @chandanpandey pointed out, sparse data structures are the way to work with this data.  I put together [this notebook](https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets/notebook) to provide some code to read the data in chunks, identify columns that have low variance (there are some columns that are all zeros), and a class to capture all those details."
  },
  "source": "meta"
}