{
  "id": 499015,
  "title": "How to handle big data in kaggle notebook",
  "url": "/competitions/leash-BELKA/discussion/499015",
  "author_name": "",
  "post_date": "2024-04-30T10:27:14.735519200Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>As you all know that size of 'Leash Bio - Predict New Medicines with BELKA' is in GBs. A Kaggle notebook cannot process such a big data. What should I do? Especially when I have a laptop with 16GB RAM only. </p>",
  "messages": [
    {
      "id": "2784523",
      "postDate": "04/30/2024 10:27:14",
      "content": "<p>As you all know that size of 'Leash Bio - Predict New Medicines with BELKA' is in GBs. A Kaggle notebook cannot process such a big data. What should I do? Especially when I have a laptop with 16GB RAM only. </p>",
      "rawMarkdown": "As you all know that size of 'Leash Bio - Predict New Medicines with BELKA' is in GBs. A Kaggle notebook cannot process such a big data. What should I do? Especially when I have a laptop with 16GB RAM only.",
      "votes": null
    },
    {
      "id": "2784556",
      "postDate": "04/30/2024 10:52:57",
      "content": "<p>I don't know how to use but there is data pipeline feature in tensorflow, which can help process data batch by batch. </p>",
      "rawMarkdown": "I don't know how to use but there is data pipeline feature in tensorflow, which can help process data batch by batch.",
      "votes": null
    },
    {
      "id": "2784577",
      "postDate": "04/30/2024 11:19:06",
      "content": "<p>TPU is an option in this case <a href=\"https://www.kaggle.com/aayushsharmacse\" target=\"_blank\">@aayushsharmacse</a> <a href=\"https://www.kaggle.com/tariqcp\" target=\"_blank\">@tariqcp</a> </p>",
      "rawMarkdown": "TPU is an option in this case @aayushsharmacse @tariqcp",
      "votes": null
    },
    {
      "id": "2784622",
      "postDate": "04/30/2024 12:00:31",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/aayushsharmacse\" target=\"_blank\">@aayushsharmacse</a> . Will TensorFlow data pipeline will handle 54GB data i.e. the training file's size in this competition.. Any link for this please. </p>",
      "rawMarkdown": "Thank you @aayushsharmacse . Will TensorFlow data pipeline will handle 54GB data i.e. the training file's size in this competition.. Any link for this please.",
      "votes": null
    },
    {
      "id": "2784626",
      "postDate": "04/30/2024 12:03:24",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Will TPU  manage 54GB data i.e. the training file's size in this competition.. Any link for this please. </p>",
      "rawMarkdown": "Thank you @ravi20076 Will TPU  manage 54GB data i.e. the training file's size in this competition.. Any link for this please.",
      "votes": null
    },
    {
      "id": "2784636",
      "postDate": "04/30/2024 12:10:34",
      "content": "<ul>\n<li>use the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491472\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/491472</a> dataset, it a compressed version</li>\n<li>use polars to open the csv faster</li>\n<li>given your limited resources, consider transforming your data (SMILES) on-the-fly into features (ECFP, embeddings …) in your DataLoader (<strong>get_item</strong> method if using Pytorch)</li>\n</ul>\n<p>I am using a 16GB machine, it works fine.</p>",
      "rawMarkdown": "use the https://www.kaggle.com/competitions/leash-BELKA/discussion/491472 dataset, it a compressed version\n- use polars to open the csv faster\n- given your limited resources, consider transforming your data (SMILES) on-the-fly into features (ECFP, embeddings ...) in your DataLoader (__get_item__ method if using Pytorch)\n\nI am using a 16GB machine, it works fine.",
      "votes": null
    },
    {
      "id": "2785412",
      "postDate": "04/30/2024 18:47:07",
      "content": "<p>Kaggle TPU gives 330GB RAM &amp; 96 CPU cores </p>",
      "rawMarkdown": "Kaggle TPU gives 330GB RAM & 96 CPU cores",
      "votes": null
    },
    {
      "id": "2788527",
      "postDate": "05/02/2024 09:04:59",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/kishanvavdara\" target=\"_blank\">@kishanvavdara</a> and obviously Kaggle as well</p>",
      "rawMarkdown": "Thank you @kishanvavdara and obviously Kaggle as well",
      "votes": null
    },
    {
      "id": "2788531",
      "postDate": "05/02/2024 09:06:19",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/louisstefanuto\" target=\"_blank\">@louisstefanuto</a>. Certainly,  it will help a lot. </p>",
      "rawMarkdown": "Thank you @louisstefanuto. Certainly,  it will help a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2784556,
      "author_name": "aayushsharmacse",
      "author_url": "",
      "post_date": "04/30/2024 10:52:57",
      "content": "<p>I don't know how to use but there is data pipeline feature in tensorflow, which can help process data batch by batch. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2784622,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "04/30/2024 12:00:31",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/aayushsharmacse\" target=\"_blank\">@aayushsharmacse</a> . Will TensorFlow data pipeline will handle 54GB data i.e. the training file's size in this competition.. Any link for this please. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2784577,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/30/2024 11:19:06",
      "content": "<p>TPU is an option in this case <a href=\"https://www.kaggle.com/aayushsharmacse\" target=\"_blank\">@aayushsharmacse</a> <a href=\"https://www.kaggle.com/tariqcp\" target=\"_blank\">@tariqcp</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2784626,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "04/30/2024 12:03:24",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Will TPU  manage 54GB data i.e. the training file's size in this competition.. Any link for this please. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2785412,
              "author_name": "kishanvavdara",
              "author_url": "",
              "post_date": "04/30/2024 18:47:07",
              "content": "<p>Kaggle TPU gives 330GB RAM &amp; 96 CPU cores </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2788527,
                  "author_name": "tariqcp",
                  "author_url": "",
                  "post_date": "05/02/2024 09:04:59",
                  "content": "<p>Thank you <a href=\"https://www.kaggle.com/kishanvavdara\" target=\"_blank\">@kishanvavdara</a> and obviously Kaggle as well</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2784636,
      "author_name": "louisstefanuto",
      "author_url": "",
      "post_date": "04/30/2024 12:10:34",
      "content": "<ul>\n<li>use the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491472\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/491472</a> dataset, it a compressed version</li>\n<li>use polars to open the csv faster</li>\n<li>given your limited resources, consider transforming your data (SMILES) on-the-fly into features (ECFP, embeddings …) in your DataLoader (<strong>get_item</strong> method if using Pytorch)</li>\n</ul>\n<p>I am using a 16GB machine, it works fine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2788531,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "05/02/2024 09:06:19",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/louisstefanuto\" target=\"_blank\">@louisstefanuto</a>. Certainly,  it will help a lot. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2784523": "As you all know that size of 'Leash Bio - Predict New Medicines with BELKA' is in GBs. A Kaggle notebook cannot process such a big data. What should I do? Especially when I have a laptop with 16GB RAM only.",
    "2784556": "I don't know how to use but there is data pipeline feature in tensorflow, which can help process data batch by batch.",
    "2784577": "TPU is an option in this case @aayushsharmacse @tariqcp",
    "2784622": "Thank you @aayushsharmacse . Will TensorFlow data pipeline will handle 54GB data i.e. the training file's size in this competition.. Any link for this please.",
    "2784626": "Thank you @ravi20076 Will TPU  manage 54GB data i.e. the training file's size in this competition.. Any link for this please.",
    "2784636": "use the https://www.kaggle.com/competitions/leash-BELKA/discussion/491472 dataset, it a compressed version\n- use polars to open the csv faster\n- given your limited resources, consider transforming your data (SMILES) on-the-fly into features (ECFP, embeddings ...) in your DataLoader (__get_item__ method if using Pytorch)\n\nI am using a 16GB machine, it works fine.",
    "2785412": "Kaggle TPU gives 330GB RAM & 96 CPU cores",
    "2788527": "Thank you @kishanvavdara and obviously Kaggle as well",
    "2788531": "Thank you @louisstefanuto. Certainly,  it will help a lot."
  },
  "source": "meta"
}