{
  "id": 363519,
  "title": "I'm so Chunked!  Better say: Trying to deal with Chunks",
  "url": "/competitions/otto-recommender-system/discussion/363519",
  "author_name": "",
  "post_date": "2022-11-02T01:36:23.513909600Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Dask  to load the data in chunks.</h1>\n<p>By Mohamed Aesawy <br>\n<a href=\"https://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77292\" target=\"_blank\">https://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77292</a></p>\n<p>import dask.dataframe as dd</p>\n<p>df = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})</p>\n<p>returns the first \"partition\" of the dataframe<br>\npart_acoustic_data = np.array(df.acoustic_data.partitions[0]) <br>\npart_time = np.array(df.time_to_failure.partitions[0])</p>\n<p>print total number of partitions<br>\nprint(df.npartitions)</p>\n<p>Spoiler alert:  I tried that and it didn't work</p>\n<p>You can try that with pandas using 'nrows' and 'chunksize' arguments of read_csv method.<br>\nComment By nroman</p>\n<h1>The data is chunked !</h1>\n<p>By Grayjay <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-nov-2021/discussion/286731\" target=\"_blank\">https://www.kaggle.com/competitions/tabular-playground-series-nov-2021/discussion/286731</a></p>\n<p>Comment by Ambros:</p>\n<p>\"We can easily visualize the chunks. If we split the data into chunks of size 20000 and plot the means of two features, we see that always three of the size-20000 chunks build a chunk of size 60000, and that the nine test chunks differ from the ten training chunks:\"</p>\n<p>train_df['chunk'] = train_df.id // 20000<br>\ntest_df['chunk'] = test_df.id // 20000<br>\nplt.scatter(train_df.groupby(['chunk']).f1.mean(), <br>\n            train_df.groupby(['chunk']).f2.mean(),<br>\n            s=3, label='train chunks')<br>\nplt.scatter(test_df.groupby(['chunk']).f1.mean(), <br>\n            test_df.groupby(['chunk']).f2.mean(), <br>\n            marker='x', label='test chunks')</p>\n<h1>TPS202112 CTGAN artifacts? Chunks?</h1>\n<p>By Rkaveland - <a href=\"https://www.kaggle.com/code/kaaveland/tps202112-ctgan-artifacts-chunks/notebook\" target=\"_blank\">https://www.kaggle.com/code/kaaveland/tps202112-ctgan-artifacts-chunks/notebook</a></p>\n<h1>Preprocessing in chunks</h1>\n<p>By Raimondo Melis - <a href=\"https://www.kaggle.com/discussions/getting-started/356189\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/356189</a><br>\nKaggle Notebook: Preprocessing of data in chunks (RIGHT WAY)<br>\n<a href=\"https://www.kaggle.com/code/raimondomelis/preprocessing-of-data-in-chunks-right-way\" target=\"_blank\">https://www.kaggle.com/code/raimondomelis/preprocessing-of-data-in-chunks-right-way</a></p>\n<h1>I hope you get chunks till the end of this Competition</h1>",
  "messages": [
    {
      "id": "2013537",
      "postDate": "11/02/2022 01:36:23",
      "content": "<h1>Dask  to load the data in chunks.</h1>\n<p>By Mohamed Aesawy <br>\n<a href=\"https://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77292\" target=\"_blank\">https://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77292</a></p>\n<p>import dask.dataframe as dd</p>\n<p>df = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})</p>\n<p>returns the first \"partition\" of the dataframe<br>\npart_acoustic_data = np.array(df.acoustic_data.partitions[0]) <br>\npart_time = np.array(df.time_to_failure.partitions[0])</p>\n<p>print total number of partitions<br>\nprint(df.npartitions)</p>\n<p>Spoiler alert:  I tried that and it didn't work</p>\n<p>You can try that with pandas using 'nrows' and 'chunksize' arguments of read_csv method.<br>\nComment By nroman</p>\n<h1>The data is chunked !</h1>\n<p>By Grayjay <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-nov-2021/discussion/286731\" target=\"_blank\">https://www.kaggle.com/competitions/tabular-playground-series-nov-2021/discussion/286731</a></p>\n<p>Comment by Ambros:</p>\n<p>\"We can easily visualize the chunks. If we split the data into chunks of size 20000 and plot the means of two features, we see that always three of the size-20000 chunks build a chunk of size 60000, and that the nine test chunks differ from the ten training chunks:\"</p>\n<p>train_df['chunk'] = train_df.id // 20000<br>\ntest_df['chunk'] = test_df.id // 20000<br>\nplt.scatter(train_df.groupby(['chunk']).f1.mean(), <br>\n            train_df.groupby(['chunk']).f2.mean(),<br>\n            s=3, label='train chunks')<br>\nplt.scatter(test_df.groupby(['chunk']).f1.mean(), <br>\n            test_df.groupby(['chunk']).f2.mean(), <br>\n            marker='x', label='test chunks')</p>\n<h1>TPS202112 CTGAN artifacts? Chunks?</h1>\n<p>By Rkaveland - <a href=\"https://www.kaggle.com/code/kaaveland/tps202112-ctgan-artifacts-chunks/notebook\" target=\"_blank\">https://www.kaggle.com/code/kaaveland/tps202112-ctgan-artifacts-chunks/notebook</a></p>\n<h1>Preprocessing in chunks</h1>\n<p>By Raimondo Melis - <a href=\"https://www.kaggle.com/discussions/getting-started/356189\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/356189</a><br>\nKaggle Notebook: Preprocessing of data in chunks (RIGHT WAY)<br>\n<a href=\"https://www.kaggle.com/code/raimondomelis/preprocessing-of-data-in-chunks-right-way\" target=\"_blank\">https://www.kaggle.com/code/raimondomelis/preprocessing-of-data-in-chunks-right-way</a></p>\n<h1>I hope you get chunks till the end of this Competition</h1>",
      "rawMarkdown": "#Dask  to load the data in chunks.\n\nBy Mohamed Aesawy \nhttps://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77292\n\nimport dask.dataframe as dd\n\ndf = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n\n returns the first \"partition\" of the dataframe\npart_acoustic_data = np.array(df.acoustic_data.partitions[0]) \npart_time = np.array(df.time_to_failure.partitions[0])\n\n print total number of partitions\nprint(df.npartitions)\n\nSpoiler alert:  I tried that and it didn't work\n\nYou can try that with pandas using 'nrows' and 'chunksize' arguments of read_csv method.\nComment By nroman\n\n#The data is chunked !\n\nBy Grayjay https://www.kaggle.com/competitions/tabular-playground-series-nov-2021/discussion/286731\n\nComment by Ambros:\n\n\"We can easily visualize the chunks. If we split the data into chunks of size 20000 and plot the means of two features, we see that always three of the size-20000 chunks build a chunk of size 60000, and that the nine test chunks differ from the ten training chunks:\"\n\ntrain_df['chunk'] = train_df.id // 20000\ntest_df['chunk'] = test_df.id // 20000\nplt.scatter(train_df.groupby(['chunk']).f1.mean(), \n            train_df.groupby(['chunk']).f2.mean(),\n            s=3, label='train chunks')\nplt.scatter(test_df.groupby(['chunk']).f1.mean(), \n            test_df.groupby(['chunk']).f2.mean(), \n            marker='x', label='test chunks')\n\n#TPS202112 CTGAN artifacts? Chunks?\n\nBy Rkaveland - https://www.kaggle.com/code/kaaveland/tps202112-ctgan-artifacts-chunks/notebook\n\n#Preprocessing in chunks\n\nBy Raimondo Melis - https://www.kaggle.com/discussions/getting-started/356189\nKaggle Notebook: Preprocessing of data in chunks (RIGHT WAY)\nhttps://www.kaggle.com/code/raimondomelis/preprocessing-of-data-in-chunks-right-way\n\n#I hope you get chunks till the end of this Competition",
      "votes": null
    },
    {
      "id": "2013763",
      "postDate": "11/02/2022 05:49:43",
      "content": "<p>Nice one <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> . Good start as always</p>",
      "rawMarkdown": "Nice one @mpwolke . Good start as always",
      "votes": null
    },
    {
      "id": "2013997",
      "postDate": "11/02/2022 08:44:04",
      "content": "<p>Great, it is very powerful.</p>",
      "rawMarkdown": "Great, it is very powerful.",
      "votes": null
    },
    {
      "id": "2014181",
      "postDate": "11/02/2022 11:49:09",
      "content": "<p>Not easy to work with them. Thank you again K. Rahman.</p>",
      "rawMarkdown": "Not easy to work with them. Thank you again K. Rahman.",
      "votes": null
    },
    {
      "id": "2014188",
      "postDate": "11/02/2022 11:58:15",
      "content": "<p>It isn't easy for beginners to work with Chunks even after reading all that users wrote about it in the previous competitions.<br>\nThank you hityangzijian.</p>",
      "rawMarkdown": "It isn't easy for beginners to work with Chunks even after reading all that users wrote about it in the previous competitions.\nThank you hityangzijian.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2013763,
      "author_name": "kalilurrahman",
      "author_url": "",
      "post_date": "11/02/2022 05:49:43",
      "content": "<p>Nice one <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> . Good start as always</p>",
      "votes": null,
      "replies": [
        {
          "id": 2014181,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "11/02/2022 11:49:09",
          "content": "<p>Not easy to work with them. Thank you again K. Rahman.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2013997,
      "author_name": "hityangzijian",
      "author_url": "",
      "post_date": "11/02/2022 08:44:04",
      "content": "<p>Great, it is very powerful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2014188,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "11/02/2022 11:58:15",
          "content": "<p>It isn't easy for beginners to work with Chunks even after reading all that users wrote about it in the previous competitions.<br>\nThank you hityangzijian.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2013537": "#Dask  to load the data in chunks.\n\nBy Mohamed Aesawy \nhttps://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77292\n\nimport dask.dataframe as dd\n\ndf = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n\n returns the first \"partition\" of the dataframe\npart_acoustic_data = np.array(df.acoustic_data.partitions[0]) \npart_time = np.array(df.time_to_failure.partitions[0])\n\n print total number of partitions\nprint(df.npartitions)\n\nSpoiler alert:  I tried that and it didn't work\n\nYou can try that with pandas using 'nrows' and 'chunksize' arguments of read_csv method.\nComment By nroman\n\n#The data is chunked !\n\nBy Grayjay https://www.kaggle.com/competitions/tabular-playground-series-nov-2021/discussion/286731\n\nComment by Ambros:\n\n\"We can easily visualize the chunks. If we split the data into chunks of size 20000 and plot the means of two features, we see that always three of the size-20000 chunks build a chunk of size 60000, and that the nine test chunks differ from the ten training chunks:\"\n\ntrain_df['chunk'] = train_df.id // 20000\ntest_df['chunk'] = test_df.id // 20000\nplt.scatter(train_df.groupby(['chunk']).f1.mean(), \n            train_df.groupby(['chunk']).f2.mean(),\n            s=3, label='train chunks')\nplt.scatter(test_df.groupby(['chunk']).f1.mean(), \n            test_df.groupby(['chunk']).f2.mean(), \n            marker='x', label='test chunks')\n\n#TPS202112 CTGAN artifacts? Chunks?\n\nBy Rkaveland - https://www.kaggle.com/code/kaaveland/tps202112-ctgan-artifacts-chunks/notebook\n\n#Preprocessing in chunks\n\nBy Raimondo Melis - https://www.kaggle.com/discussions/getting-started/356189\nKaggle Notebook: Preprocessing of data in chunks (RIGHT WAY)\nhttps://www.kaggle.com/code/raimondomelis/preprocessing-of-data-in-chunks-right-way\n\n#I hope you get chunks till the end of this Competition",
    "2013763": "Nice one @mpwolke . Good start as always",
    "2013997": "Great, it is very powerful.",
    "2014181": "Not easy to work with them. Thank you again K. Rahman.",
    "2014188": "It isn't easy for beginners to work with Chunks even after reading all that users wrote about it in the previous competitions.\nThank you hityangzijian."
  },
  "source": "meta"
}