{
  "id": 348827,
  "title": "Need help to handle huge datasets.",
  "url": "/competitions/open-problems-multimodal/discussion/348827",
  "author_name": "",
  "post_date": "2022-08-30T06:49:39.594758900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Kagglers, recently I started my data science journey in Kaggle, I always wanted to participate in competitions and use my knowledge to try to get into the game. But I am stuck and confused about handling large datasets. I have a laptop with the specs 1650 GTX with 16GB RAM and the model name is Asus ROG STRIX G17. How can I load huge datasets without killing my ram? I am trying to find answers to get started in competitions and also start practising on really huge datasets. Any suggestions?</p>",
  "messages": [
    {
      "id": "1919134",
      "postDate": "08/30/2022 06:49:39",
      "content": "<p>Hi Kagglers, recently I started my data science journey in Kaggle, I always wanted to participate in competitions and use my knowledge to try to get into the game. But I am stuck and confused about handling large datasets. I have a laptop with the specs 1650 GTX with 16GB RAM and the model name is Asus ROG STRIX G17. How can I load huge datasets without killing my ram? I am trying to find answers to get started in competitions and also start practising on really huge datasets. Any suggestions?</p>",
      "rawMarkdown": "Hi Kagglers, recently I started my data science journey in Kaggle, I always wanted to participate in competitions and use my knowledge to try to get into the game. But I am stuck and confused about handling large datasets. I have a laptop with the specs 1650 GTX with 16GB RAM and the model name is Asus ROG STRIX G17. How can I load huge datasets without killing my ram? I am trying to find answers to get started in competitions and also start practising on really huge datasets. Any suggestions?",
      "votes": null
    },
    {
      "id": "1919472",
      "postDate": "08/30/2022 12:45:59",
      "content": "<p>Ok specifically for competitions, I usually load the data into the kaggle directory and use Kaggle notebook for the competition. For this competition specifically, it would look something like:<br>\n<a href=\"https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading\" target=\"_blank\">https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading</a></p>\n<p>However, for working with large datasets there are some data processing systems that help with big data workloads (I haven't used them yet but I plan on checking them out. I would definitely ask someone about them if you know someone personally) I've heard of Spark and Hadoop (or PySpark maybe)</p>\n<p>I hope this helps, good luck with the competition, and happy kaggling!!</p>",
      "rawMarkdown": "Ok specifically for competitions, I usually load the data into the kaggle directory and use Kaggle notebook for the competition. For this competition specifically, it would look something like:\nhttps://www.kaggle.com/code/peterholderrieth/getting-started-data-loading\n\nHowever, for working with large datasets there are some data processing systems that help with big data workloads (I haven't used them yet but I plan on checking them out. I would definitely ask someone about them if you know someone personally) I've heard of Spark and Hadoop (or PySpark maybe)\n\nI hope this helps, good luck with the competition, and happy kaggling!!",
      "votes": null
    },
    {
      "id": "1919804",
      "postDate": "08/30/2022 17:20:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pybeast\" target=\"_blank\">@pybeast</a>, I never thought of loading the dataset into the Kaggle notebook (I will try this). And also, I have heard about PyTorch, I should definitely try it. Thanks you :D</p>",
      "rawMarkdown": "Hi @pybeast, I never thought of loading the dataset into the Kaggle notebook (I will try this). And also, I have heard about PyTorch, I should definitely try it. Thanks you :D",
      "votes": null
    },
    {
      "id": "1920048",
      "postDate": "08/30/2022 20:24:55",
      "content": "<p>Hi Neeresh,</p>\n<p>the competition organizers have arranged an agreement with a cloud computing provider, that gives access to large computing capabilities (machines with 128 Gb of RAM, 16 Gb GPU, 40 Gb of storage) to participants. Check out this discussion: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999</a>.</p>\n<p>Besides that, there are a few discussions about this. One I've started: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690</a>, and a notebook about using csr matrices (since much of the data is really sparse, especially multiome inputs): <a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices\" target=\"_blank\">https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices</a>.</p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hi Neeresh,\n\nthe competition organizers have arranged an agreement with a cloud computing provider, that gives access to large computing capabilities (machines with 128 Gb of RAM, 16 Gb GPU, 40 Gb of storage) to participants. Check out this discussion: [https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999).\n\nBesides that, there are a few discussions about this. One I've started: [https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690), and a notebook about using csr matrices (since much of the data is really sparse, especially multiome inputs): [https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices](https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices).\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "1920102",
      "postDate": "08/30/2022 22:12:23",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/neereshkumar\" target=\"_blank\">@neereshkumar</a>. If dimensionality reduction is not an option then maybe you could look for some virtual machine at some cloud service provider. I have already used (not for kaggle competitions yet) virtual machines (EC2) at AWS (Amazon). Costxbenefit seemed ok to me. </p>",
      "rawMarkdown": "Hello @neereshkumar. If dimensionality reduction is not an option then maybe you could look for some virtual machine at some cloud service provider. I have already used (not for kaggle competitions yet) virtual machines (EC2) at AWS (Amazon). Costxbenefit seemed ok to me.",
      "votes": null
    },
    {
      "id": "1920350",
      "postDate": "08/31/2022 05:19:59",
      "content": "<p>oh… this is new to me. Thank you for sharing it. I will check.</p>",
      "rawMarkdown": "oh... this is new to me. Thank you for sharing it. I will check.",
      "votes": null
    },
    {
      "id": "1920352",
      "postDate": "08/31/2022 05:21:45",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/clacflores\" target=\"_blank\">@clacflores</a>, yes right now I am using Azure, but I need a way to handle large datasets without using cloud services.</p>",
      "rawMarkdown": "Hi @clacflores, yes right now I am using Azure, but I need a way to handle large datasets without using cloud services.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1919472,
      "author_name": "pybeast",
      "author_url": "",
      "post_date": "08/30/2022 12:45:59",
      "content": "<p>Ok specifically for competitions, I usually load the data into the kaggle directory and use Kaggle notebook for the competition. For this competition specifically, it would look something like:<br>\n<a href=\"https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading\" target=\"_blank\">https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading</a></p>\n<p>However, for working with large datasets there are some data processing systems that help with big data workloads (I haven't used them yet but I plan on checking them out. I would definitely ask someone about them if you know someone personally) I've heard of Spark and Hadoop (or PySpark maybe)</p>\n<p>I hope this helps, good luck with the competition, and happy kaggling!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1919804,
          "author_name": "neereshkumar",
          "author_url": "",
          "post_date": "08/30/2022 17:20:25",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pybeast\" target=\"_blank\">@pybeast</a>, I never thought of loading the dataset into the Kaggle notebook (I will try this). And also, I have heard about PyTorch, I should definitely try it. Thanks you :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1920048,
      "author_name": "alekeuro",
      "author_url": "",
      "post_date": "08/30/2022 20:24:55",
      "content": "<p>Hi Neeresh,</p>\n<p>the competition organizers have arranged an agreement with a cloud computing provider, that gives access to large computing capabilities (machines with 128 Gb of RAM, 16 Gb GPU, 40 Gb of storage) to participants. Check out this discussion: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999</a>.</p>\n<p>Besides that, there are a few discussions about this. One I've started: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690</a>, and a notebook about using csr matrices (since much of the data is really sparse, especially multiome inputs): <a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices\" target=\"_blank\">https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices</a>.</p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1920350,
          "author_name": "neereshkumar",
          "author_url": "",
          "post_date": "08/31/2022 05:19:59",
          "content": "<p>oh… this is new to me. Thank you for sharing it. I will check.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1920102,
      "author_name": "clacflores",
      "author_url": "",
      "post_date": "08/30/2022 22:12:23",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/neereshkumar\" target=\"_blank\">@neereshkumar</a>. If dimensionality reduction is not an option then maybe you could look for some virtual machine at some cloud service provider. I have already used (not for kaggle competitions yet) virtual machines (EC2) at AWS (Amazon). Costxbenefit seemed ok to me. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1920352,
          "author_name": "neereshkumar",
          "author_url": "",
          "post_date": "08/31/2022 05:21:45",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/clacflores\" target=\"_blank\">@clacflores</a>, yes right now I am using Azure, but I need a way to handle large datasets without using cloud services.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1919134": "Hi Kagglers, recently I started my data science journey in Kaggle, I always wanted to participate in competitions and use my knowledge to try to get into the game. But I am stuck and confused about handling large datasets. I have a laptop with the specs 1650 GTX with 16GB RAM and the model name is Asus ROG STRIX G17. How can I load huge datasets without killing my ram? I am trying to find answers to get started in competitions and also start practising on really huge datasets. Any suggestions?",
    "1919472": "Ok specifically for competitions, I usually load the data into the kaggle directory and use Kaggle notebook for the competition. For this competition specifically, it would look something like:\nhttps://www.kaggle.com/code/peterholderrieth/getting-started-data-loading\n\nHowever, for working with large datasets there are some data processing systems that help with big data workloads (I haven't used them yet but I plan on checking them out. I would definitely ask someone about them if you know someone personally) I've heard of Spark and Hadoop (or PySpark maybe)\n\nI hope this helps, good luck with the competition, and happy kaggling!!",
    "1919804": "Hi @pybeast, I never thought of loading the dataset into the Kaggle notebook (I will try this). And also, I have heard about PyTorch, I should definitely try it. Thanks you :D",
    "1920048": "Hi Neeresh,\n\nthe competition organizers have arranged an agreement with a cloud computing provider, that gives access to large computing capabilities (machines with 128 Gb of RAM, 16 Gb GPU, 40 Gb of storage) to participants. Check out this discussion: [https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999).\n\nBesides that, there are a few discussions about this. One I've started: [https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690), and a notebook about using csr matrices (since much of the data is really sparse, especially multiome inputs): [https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices](https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices).\n\nHope this helps!",
    "1920102": "Hello @neereshkumar. If dimensionality reduction is not an option then maybe you could look for some virtual machine at some cloud service provider. I have already used (not for kaggle competitions yet) virtual machines (EC2) at AWS (Amazon). Costxbenefit seemed ok to me.",
    "1920350": "oh... this is new to me. Thank you for sharing it. I will check.",
    "1920352": "Hi @clacflores, yes right now I am using Azure, but I need a way to handle large datasets without using cloud services."
  },
  "source": "meta"
}