{
  "id": 327084,
  "title": "Is this the largest tabular dataset on Kaggle?",
  "url": "/competitions/amex-default-prediction/discussion/327084",
  "author_name": "",
  "post_date": "2022-05-25T14:32:14.582295400Z",
  "votes": 32,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I guess the notebooks won't be relevant here. If we use numpy, we might barely fit the training data into 16 GBs of memory.</p>",
  "messages": [
    {
      "id": "1801226",
      "postDate": "05/25/2022 14:32:14",
      "content": "<p>I guess the notebooks won't be relevant here. If we use numpy, we might barely fit the training data into 16 GBs of memory.</p>",
      "rawMarkdown": "I guess the notebooks won't be relevant here. If we use numpy, we might barely fit the training data into 16 GBs of memory.",
      "votes": null
    },
    {
      "id": "1801264",
      "postDate": "05/25/2022 15:08:36",
      "content": "<p>Personal computer memory is not enough to drive huge datasets.</p>",
      "rawMarkdown": "Personal computer memory is not enough to drive huge datasets.",
      "votes": null
    },
    {
      "id": "1801289",
      "postDate": "05/25/2022 15:39:07",
      "content": "<p>Nope its not the biggest one.</p>\n<p>Here is the link for even bigger dataset (and \"raw\" data).<br>\n<a href=\"https://www.kaggle.com/competitions/avito-demand-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/avito-demand-prediction/data</a></p>",
      "rawMarkdown": "Nope its not the biggest one.\n\nHere is the link for even bigger dataset (and \"raw\" data).\nhttps://www.kaggle.com/competitions/avito-demand-prediction/data",
      "votes": null
    },
    {
      "id": "1801442",
      "postDate": "05/25/2022 18:36:53",
      "content": "<p>But that one has image data.</p>",
      "rawMarkdown": "But that one has image data.",
      "votes": null
    },
    {
      "id": "1801453",
      "postDate": "05/25/2022 19:08:38",
      "content": "<p>Only winners will be paid back for electrical bills 😁</p>",
      "rawMarkdown": "Only winners will be paid back for electrical bills 😁",
      "votes": null
    },
    {
      "id": "1801885",
      "postDate": "05/26/2022 08:36:16",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> you are totally right - only tabular data there  ~25Gb<br>\nI do remember a forecasting competition with may sql files like this one:<br>\n<a href=\"https://www.kaggle.com/competitions/avito-context-ad-clicks/data\" target=\"_blank\">https://www.kaggle.com/competitions/avito-context-ad-clicks/data</a><br>\nbut not exactly this.</p>\n<p>My memory is telling that there were multiple raw files with 100Gb+ of pure tabular data. I can be wrong.</p>",
      "rawMarkdown": "philippsinger you are totally right - only tabular data there  ~25Gb\nI do remember a forecasting competition with may sql files like this one:\nhttps://www.kaggle.com/competitions/avito-context-ad-clicks/data\nbut not exactly this.\n\nMy memory is telling that there were multiple raw files with 100Gb+ of pure tabular data. I can be wrong.",
      "votes": null
    },
    {
      "id": "1802527",
      "postDate": "05/26/2022 20:26:33",
      "content": "<p>Sampling will be your friend</p>",
      "rawMarkdown": "Sampling will be your friend",
      "votes": null
    },
    {
      "id": "1802528",
      "postDate": "05/26/2022 20:27:25",
      "content": "<p>Half of the success lies in how to deal with it😏</p>",
      "rawMarkdown": "Half of the success lies in how to deal with it😏",
      "votes": null
    },
    {
      "id": "1802541",
      "postDate": "05/26/2022 21:00:14",
      "content": "<p>I have had great success with Personal Computer memory on Ubuntu using a very large swap file (200-300GB) on a SSD drive - with 64GB of hardware ram.</p>",
      "rawMarkdown": "I have had great success with Personal Computer memory on Ubuntu using a very large swap file (200-300GB) on a SSD drive - with 64GB of hardware ram.",
      "votes": null
    },
    {
      "id": "1803257",
      "postDate": "05/27/2022 15:59:51",
      "content": "<p>I'm using an Ubuntu server with 2 RTX 3060s and 32 GB of RAM, hopefully that should be enough.</p>",
      "rawMarkdown": "I'm using an Ubuntu server with 2 RTX 3060s and 32 GB of RAM, hopefully that should be enough.",
      "votes": null
    },
    {
      "id": "1803560",
      "postDate": "05/28/2022 00:12:10",
      "content": "<p>I believe Colab Pro+ should be able to handle it on a high ram machine; I will give it a try. also, the feathers datasets going around are pretty helpful, and I have never seen a competition with so much data, so probable it's the biggest Kaggle Tabular Dataset.</p>",
      "rawMarkdown": "I believe Colab Pro+ should be able to handle it on a high ram machine; I will give it a try. also, the feathers datasets going around are pretty helpful, and I have never seen a competition with so much data, so probable it's the biggest Kaggle Tabular Dataset.",
      "votes": null
    },
    {
      "id": "1803822",
      "postDate": "05/28/2022 07:56:34",
      "content": "<p>Yep, looks like it will be enough. I'm also unleashing my secret weapon. I just downloaded the data on my companies workstation with AMD Ryzen 9 5950X 16-Core Processor, NVIDIA GeForce RTX 3090 and 64 GB RAM.</p>",
      "rawMarkdown": "Yep, looks like it will be enough. I'm also unleashing my secret weapon. I just downloaded the data on my companies workstation with AMD Ryzen 9 5950X 16-Core Processor, NVIDIA GeForce RTX 3090 and 64 GB RAM.",
      "votes": null
    },
    {
      "id": "1804531",
      "postDate": "05/29/2022 05:51:20",
      "content": "<p>Hope you get a good grade.👊</p>",
      "rawMarkdown": "Hope you get a good grade.👊",
      "votes": null
    },
    {
      "id": "1804841",
      "postDate": "05/29/2022 13:50:39",
      "content": "<p>5 million rows training data and more than 10 million rows test data.  That's really a big dataset. But I managed to train and inference on Kaggle Notebook.</p>",
      "rawMarkdown": "5 million rows training data and more than 10 million rows test data.  That's really a big dataset. But I managed to train and inference on Kaggle Notebook.",
      "votes": null
    },
    {
      "id": "1805075",
      "postDate": "05/29/2022 18:57:00",
      "content": "<p>testing this on my windows / RTX 2070 / 64 Gb Ram…</p>",
      "rawMarkdown": "testing this on my windows / RTX 2070 / 64 Gb Ram...",
      "votes": null
    },
    {
      "id": "1806635",
      "postDate": "05/31/2022 10:54:57",
      "content": "<p>We need to move to Keban. :D</p>",
      "rawMarkdown": "We need to move to Keban. :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801264,
      "author_name": "kgxiao",
      "author_url": "",
      "post_date": "05/25/2022 15:08:36",
      "content": "<p>Personal computer memory is not enough to drive huge datasets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802527,
          "author_name": "michaelburnes",
          "author_url": "",
          "post_date": "05/26/2022 20:26:33",
          "content": "<p>Sampling will be your friend</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802541,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "05/26/2022 21:00:14",
          "content": "<p>I have had great success with Personal Computer memory on Ubuntu using a very large swap file (200-300GB) on a SSD drive - with 64GB of hardware ram.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1804531,
          "author_name": "kgxiao",
          "author_url": "",
          "post_date": "05/29/2022 05:51:20",
          "content": "<p>Hope you get a good grade.👊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801289,
      "author_name": "kyakovlev",
      "author_url": "",
      "post_date": "05/25/2022 15:39:07",
      "content": "<p>Nope its not the biggest one.</p>\n<p>Here is the link for even bigger dataset (and \"raw\" data).<br>\n<a href=\"https://www.kaggle.com/competitions/avito-demand-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/avito-demand-prediction/data</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1801442,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/25/2022 18:36:53",
          "content": "<p>But that one has image data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801885,
          "author_name": "kyakovlev",
          "author_url": "",
          "post_date": "05/26/2022 08:36:16",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> you are totally right - only tabular data there  ~25Gb<br>\nI do remember a forecasting competition with may sql files like this one:<br>\n<a href=\"https://www.kaggle.com/competitions/avito-context-ad-clicks/data\" target=\"_blank\">https://www.kaggle.com/competitions/avito-context-ad-clicks/data</a><br>\nbut not exactly this.</p>\n<p>My memory is telling that there were multiple raw files with 100Gb+ of pure tabular data. I can be wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801453,
      "author_name": "huseyincot",
      "author_url": "",
      "post_date": "05/25/2022 19:08:38",
      "content": "<p>Only winners will be paid back for electrical bills 😁</p>",
      "votes": null,
      "replies": [
        {
          "id": 1806635,
          "author_name": "mozturkmen",
          "author_url": "",
          "post_date": "05/31/2022 10:54:57",
          "content": "<p>We need to move to Keban. :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802528,
      "author_name": "katearb",
      "author_url": "",
      "post_date": "05/26/2022 20:27:25",
      "content": "<p>Half of the success lies in how to deal with it😏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1803257,
      "author_name": "sbhatti",
      "author_url": "",
      "post_date": "05/27/2022 15:59:51",
      "content": "<p>I'm using an Ubuntu server with 2 RTX 3060s and 32 GB of RAM, hopefully that should be enough.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1803822,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "05/28/2022 07:56:34",
          "content": "<p>Yep, looks like it will be enough. I'm also unleashing my secret weapon. I just downloaded the data on my companies workstation with AMD Ryzen 9 5950X 16-Core Processor, NVIDIA GeForce RTX 3090 and 64 GB RAM.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1803560,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "05/28/2022 00:12:10",
      "content": "<p>I believe Colab Pro+ should be able to handle it on a high ram machine; I will give it a try. also, the feathers datasets going around are pretty helpful, and I have never seen a competition with so much data, so probable it's the biggest Kaggle Tabular Dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1804841,
      "author_name": "lonnieqin",
      "author_url": "",
      "post_date": "05/29/2022 13:50:39",
      "content": "<p>5 million rows training data and more than 10 million rows test data.  That's really a big dataset. But I managed to train and inference on Kaggle Notebook.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1805075,
      "author_name": "drzamora",
      "author_url": "",
      "post_date": "05/29/2022 18:57:00",
      "content": "<p>testing this on my windows / RTX 2070 / 64 Gb Ram…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801226": "I guess the notebooks won't be relevant here. If we use numpy, we might barely fit the training data into 16 GBs of memory.",
    "1801264": "Personal computer memory is not enough to drive huge datasets.",
    "1801289": "Nope its not the biggest one.\n\nHere is the link for even bigger dataset (and \"raw\" data).\nhttps://www.kaggle.com/competitions/avito-demand-prediction/data",
    "1801442": "But that one has image data.",
    "1801453": "Only winners will be paid back for electrical bills 😁",
    "1801885": "philippsinger you are totally right - only tabular data there  ~25Gb\nI do remember a forecasting competition with may sql files like this one:\nhttps://www.kaggle.com/competitions/avito-context-ad-clicks/data\nbut not exactly this.\n\nMy memory is telling that there were multiple raw files with 100Gb+ of pure tabular data. I can be wrong.",
    "1802527": "Sampling will be your friend",
    "1802528": "Half of the success lies in how to deal with it😏",
    "1802541": "I have had great success with Personal Computer memory on Ubuntu using a very large swap file (200-300GB) on a SSD drive - with 64GB of hardware ram.",
    "1803257": "I'm using an Ubuntu server with 2 RTX 3060s and 32 GB of RAM, hopefully that should be enough.",
    "1803560": "I believe Colab Pro+ should be able to handle it on a high ram machine; I will give it a try. also, the feathers datasets going around are pretty helpful, and I have never seen a competition with so much data, so probable it's the biggest Kaggle Tabular Dataset.",
    "1803822": "Yep, looks like it will be enough. I'm also unleashing my secret weapon. I just downloaded the data on my companies workstation with AMD Ryzen 9 5950X 16-Core Processor, NVIDIA GeForce RTX 3090 and 64 GB RAM.",
    "1804531": "Hope you get a good grade.👊",
    "1804841": "5 million rows training data and more than 10 million rows test data.  That's really a big dataset. But I managed to train and inference on Kaggle Notebook.",
    "1805075": "testing this on my windows / RTX 2070 / 64 Gb Ram...",
    "1806635": "We need to move to Keban. :D"
  },
  "source": "meta"
}