{
  "id": 392044,
  "title": "How do you handle such a huge dataset??",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/392044",
  "author_name": "",
  "post_date": "2023-03-03T13:57:48.320137400Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I tried to apply aggregation to all events, but too many hous passed.<br>\nDo you have any ideas to handle it?</p>\n<p>BigQuery? or cuDF?</p>\n<p>I'm very bigginer, so I hope your advice!<br>\nThanks in advance.</p>",
  "messages": [
    {
      "id": "2167457",
      "postDate": "03/03/2023 13:57:48",
      "content": "<p>I tried to apply aggregation to all events, but too many hous passed.<br>\nDo you have any ideas to handle it?</p>\n<p>BigQuery? or cuDF?</p>\n<p>I'm very bigginer, so I hope your advice!<br>\nThanks in advance.</p>",
      "rawMarkdown": "I tried to apply aggregation to all events, but too many hous passed.\nDo you have any ideas to handle it?\n\nBigQuery? or cuDF?\n\nI'm very bigginer, so I hope your advice!\nThanks in advance.",
      "votes": null
    },
    {
      "id": "2168206",
      "postDate": "03/04/2023 02:08:27",
      "content": "<p>A few ideas: <br>\n1) Don't combine all batches, just process them subsequently<br>\n2) Select a subset of batches and/or events for efficient testing<br>\n3) Split the event data (train/test_meta.parquet) into batches</p>",
      "rawMarkdown": "A few ideas: \n1) Don't combine all batches, just process them subsequently\n2) Select a subset of batches and/or events for efficient testing\n3) Split the event data (train/test_meta.parquet) into batches",
      "votes": null
    },
    {
      "id": "2168342",
      "postDate": "03/04/2023 06:19:33",
      "content": "<p>Thanks! I'll take them into account!</p>",
      "rawMarkdown": "Thanks! I'll take them into account!",
      "votes": null
    },
    {
      "id": "2171134",
      "postDate": "03/06/2023 14:39:28",
      "content": "<p>buy more memory modules😂 , as resently they are cheap（just kidding）</p>",
      "rawMarkdown": "buy more memory modules😂 , as resently they are cheap（just kidding）",
      "votes": null
    },
    {
      "id": "2171139",
      "postDate": "03/06/2023 14:41:30",
      "content": "<p>you can read notebook like this <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">notebook</a>, the contest also have big dataset</p>",
      "rawMarkdown": "you can read notebook like this [notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575), the contest also have big dataset",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168206,
      "author_name": "taqseorangpun",
      "author_url": "",
      "post_date": "03/04/2023 02:08:27",
      "content": "<p>A few ideas: <br>\n1) Don't combine all batches, just process them subsequently<br>\n2) Select a subset of batches and/or events for efficient testing<br>\n3) Split the event data (train/test_meta.parquet) into batches</p>",
      "votes": null,
      "replies": [
        {
          "id": 2168342,
          "author_name": "daichiuraie",
          "author_url": "",
          "post_date": "03/04/2023 06:19:33",
          "content": "<p>Thanks! I'll take them into account!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2171134,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "03/06/2023 14:39:28",
      "content": "<p>buy more memory modules😂 , as resently they are cheap（just kidding）</p>",
      "votes": null,
      "replies": [
        {
          "id": 2171139,
          "author_name": "roger92",
          "author_url": "",
          "post_date": "03/06/2023 14:41:30",
          "content": "<p>you can read notebook like this <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">notebook</a>, the contest also have big dataset</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2167457": "I tried to apply aggregation to all events, but too many hous passed.\nDo you have any ideas to handle it?\n\nBigQuery? or cuDF?\n\nI'm very bigginer, so I hope your advice!\nThanks in advance.",
    "2168206": "A few ideas: \n1) Don't combine all batches, just process them subsequently\n2) Select a subset of batches and/or events for efficient testing\n3) Split the event data (train/test_meta.parquet) into batches",
    "2168342": "Thanks! I'll take them into account!",
    "2171134": "buy more memory modules😂 , as resently they are cheap（just kidding）",
    "2171139": "you can read notebook like this [notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575), the contest also have big dataset"
  },
  "source": "meta"
}