{
  "id": 491328,
  "title": "The train data has 295,246,830 rows !!",
  "url": "/competitions/leash-BELKA/discussion/491328",
  "author_name": "",
  "post_date": "2024-04-05T12:57:29.171143500Z",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I opened the train data with polars in an environment with sufficient memory. <br>\nThe train data has a shape of (295,246,830, 7).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fca53e3213d54dec1a5cf7de39df7e15c%2FClipboard01.jpg?generation=1712321784076111&amp;alt=media\"></p>\n<p>Using polars, I can open the parquet file in 17 seconds with about 70Gb memory.</p>\n<p>※ On the other hand, the test data has a shape of (1,674,896, 6).</p>\n<p>Have there been competitions in the past with training data as large as this? Would it be necessary to use all the data to win?</p>",
  "messages": [
    {
      "id": "2736879",
      "postDate": "04/05/2024 12:57:29",
      "content": "<p>I opened the train data with polars in an environment with sufficient memory. <br>\nThe train data has a shape of (295,246,830, 7).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fca53e3213d54dec1a5cf7de39df7e15c%2FClipboard01.jpg?generation=1712321784076111&amp;alt=media\"></p>\n<p>Using polars, I can open the parquet file in 17 seconds with about 70Gb memory.</p>\n<p>※ On the other hand, the test data has a shape of (1,674,896, 6).</p>\n<p>Have there been competitions in the past with training data as large as this? Would it be necessary to use all the data to win?</p>",
      "rawMarkdown": "I opened the train data with polars in an environment with sufficient memory. \nThe train data has a shape of (295,246,830, 7).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fca53e3213d54dec1a5cf7de39df7e15c%2FClipboard01.jpg?generation=1712321784076111&alt=media)\n\nUsing polars, I can open the parquet file in 17 seconds with about 70Gb memory.\n\n※ On the other hand, the test data has a shape of (1,674,896, 6).\n\nHave there been competitions in the past with training data as large as this? Would it be necessary to use all the data to win?",
      "votes": null
    },
    {
      "id": "2736890",
      "postDate": "04/05/2024 13:09:38",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">stanford-ribonanza</a> competition had somewhat similar train size. </p>",
      "rawMarkdown": "[stanford-ribonanza](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data) competition had somewhat similar train size.",
      "votes": null
    },
    {
      "id": "2736914",
      "postDate": "04/05/2024 13:26:14",
      "content": "<p><a href=\"https://www.kaggle.com/kishanvavdara\" target=\"_blank\">@kishanvavdara</a> Thank you for telling me! I will check the solution for stanford-ribonanza!!</p>",
      "rawMarkdown": "kishanvavdara Thank you for telling me! I will check the solution for stanford-ribonanza!!",
      "votes": null
    },
    {
      "id": "2737051",
      "postDate": "04/05/2024 14:47:41",
      "content": "<p>Well, kind of, if we count by labels since Ribonanza had a label per nucleotide, but only ~1M samples (so 1M samples times ~200 labels per sample = ~200M).<br>\nBy sheer data size, <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">IceCube</a> dwarfs both with ~117GB.</p>",
      "rawMarkdown": "Well, kind of, if we count by labels since Ribonanza had a label per nucleotide, but only ~1M samples (so 1M samples times ~200 labels per sample = ~200M).\nBy sheer data size, [IceCube](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) dwarfs both with ~117GB.",
      "votes": null
    },
    {
      "id": "2737090",
      "postDate": "04/05/2024 15:19:44",
      "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> Thank you for comment ! I'll check it!</p>",
      "rawMarkdown": "shlomoron Thank you for comment ! I'll check it!",
      "votes": null
    },
    {
      "id": "2738114",
      "postDate": "04/06/2024 06:02:22",
      "content": "<p>train data has 295,246,830 rows !!       how this speed to run and time consumed in the environment??</p>",
      "rawMarkdown": "train data has 295,246,830 rows !!       how this speed to run and time consumed in the environment??",
      "votes": null
    },
    {
      "id": "2738600",
      "postDate": "04/06/2024 12:48:56",
      "content": "<p>Ribonanza had around ~1M samples (in terms of computational complexity, since for each sample we got 200-400 labels), this competition has 300M<br>\nNot sure if we can do that without downsample :D </p>",
      "rawMarkdown": "Ribonanza had around ~1M samples (in terms of computational complexity, since for each sample we got 200-400 labels), this competition has 300M\nNot sure if we can do that without downsample :D",
      "votes": null
    },
    {
      "id": "2738724",
      "postDate": "04/06/2024 14:46:31",
      "content": "<p>Just feed the TPU a batch of 1k, np :)<br>\np.s. It's actually 100M with three labels per sample.</p>",
      "rawMarkdown": "Just feed the TPU a batch of 1k, np :)\np.s. It's actually 100M with three labels per sample.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2736890,
      "author_name": "kishanvavdara",
      "author_url": "",
      "post_date": "04/05/2024 13:09:38",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">stanford-ribonanza</a> competition had somewhat similar train size. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2736914,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "04/05/2024 13:26:14",
          "content": "<p><a href=\"https://www.kaggle.com/kishanvavdara\" target=\"_blank\">@kishanvavdara</a> Thank you for telling me! I will check the solution for stanford-ribonanza!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2737051,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "04/05/2024 14:47:41",
          "content": "<p>Well, kind of, if we count by labels since Ribonanza had a label per nucleotide, but only ~1M samples (so 1M samples times ~200 labels per sample = ~200M).<br>\nBy sheer data size, <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">IceCube</a> dwarfs both with ~117GB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2737090,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "04/05/2024 15:19:44",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> Thank you for comment ! I'll check it!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2738114,
      "author_name": "sonialikhan",
      "author_url": "",
      "post_date": "04/06/2024 06:02:22",
      "content": "<p>train data has 295,246,830 rows !!       how this speed to run and time consumed in the environment??</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2738600,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "04/06/2024 12:48:56",
      "content": "<p>Ribonanza had around ~1M samples (in terms of computational complexity, since for each sample we got 200-400 labels), this competition has 300M<br>\nNot sure if we can do that without downsample :D </p>",
      "votes": null,
      "replies": [
        {
          "id": 2738724,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "04/06/2024 14:46:31",
          "content": "<p>Just feed the TPU a batch of 1k, np :)<br>\np.s. It's actually 100M with three labels per sample.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2736879": "I opened the train data with polars in an environment with sufficient memory. \nThe train data has a shape of (295,246,830, 7).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fca53e3213d54dec1a5cf7de39df7e15c%2FClipboard01.jpg?generation=1712321784076111&alt=media)\n\nUsing polars, I can open the parquet file in 17 seconds with about 70Gb memory.\n\n※ On the other hand, the test data has a shape of (1,674,896, 6).\n\nHave there been competitions in the past with training data as large as this? Would it be necessary to use all the data to win?",
    "2736890": "[stanford-ribonanza](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data) competition had somewhat similar train size.",
    "2736914": "kishanvavdara Thank you for telling me! I will check the solution for stanford-ribonanza!!",
    "2737051": "Well, kind of, if we count by labels since Ribonanza had a label per nucleotide, but only ~1M samples (so 1M samples times ~200 labels per sample = ~200M).\nBy sheer data size, [IceCube](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) dwarfs both with ~117GB.",
    "2737090": "shlomoron Thank you for comment ! I'll check it!",
    "2738114": "train data has 295,246,830 rows !!       how this speed to run and time consumed in the environment??",
    "2738600": "Ribonanza had around ~1M samples (in terms of computational complexity, since for each sample we got 200-400 labels), this competition has 300M\nNot sure if we can do that without downsample :D",
    "2738724": "Just feed the TPU a batch of 1k, np :)\np.s. It's actually 100M with three labels per sample."
  },
  "source": "meta"
}