{
  "id": 363843,
  "title": "Full dataset processed to CSV/parquet files with optimized memory footprint",
  "url": "/competitions/otto-recommender-system/discussion/363843",
  "author_name": "",
  "post_date": "2022-11-03T11:25:01.421683100Z",
  "votes": 77,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Hey friends!</p>\n<p>I converted the dataset to live in <code>csv</code> &amp; <code>parquet</code> files. You can find the dataset <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">here</a>.</p>\n<p>I have also created <a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-file\" target=\"_blank\">an explainer notebook</a>. In it, I run you through how to use this data.</p>\n<p>I am also attaching the code I used to process this data as a notebook file to this post. I couldn't process the data on Kaggle due to RAM limitations so I ran it on my local machine and uploaded it to Kaggle.</p>\n<p>Hope this can speed you along in your work 🙂 Happy kaggling! 🥳</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2015554",
      "postDate": "11/03/2022 11:25:01",
      "content": "<p>Hey friends!</p>\n<p>I converted the dataset to live in <code>csv</code> &amp; <code>parquet</code> files. You can find the dataset <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">here</a>.</p>\n<p>I have also created <a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-file\" target=\"_blank\">an explainer notebook</a>. In it, I run you through how to use this data.</p>\n<p>I am also attaching the code I used to process this data as a notebook file to this post. I couldn't process the data on Kaggle due to RAM limitations so I ran it on my local machine and uploaded it to Kaggle.</p>\n<p>Hope this can speed you along in your work 🙂 Happy kaggling! 🥳</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey friends!\n\nI converted the dataset to live in `csv` & `parquet` files. You can find the dataset [here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).\n\nI have also created [an explainer notebook](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-file). In it, I run you through how to use this data.\n\nI am also attaching the code I used to process this data as a notebook file to this post. I couldn't process the data on Kaggle due to RAM limitations so I ran it on my local machine and uploaded it to Kaggle.\n\nHope this can speed you along in your work 🙂 Happy kaggling! 🥳\n\n\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2015604",
      "postDate": "11/03/2022 12:16:51",
      "content": "<p>This is the script I used for preprocessing as a screenshot</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fd1438821834f8779624042ba02b4d132%2Fpreprocessing_script.png?generation=1667477805785949&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This is the script I used for preprocessing as a screenshot\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fd1438821834f8779624042ba02b4d132%2Fpreprocessing_script.png?generation=1667477805785949&alt=media)",
      "votes": null
    },
    {
      "id": "2015683",
      "postDate": "11/03/2022 13:05:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ! Glad to see you here.</p>\n<p>I was taking a look at the data and I have seen that for example <strong>session 0</strong> in train data lasts from <strong>July 31 (ts → 1659304800025)</strong> to <strong>August 28 (ts → 1661684983707)</strong></p>\n<p>In my concept of a session, lasting a month seems too much, I don't know what you think. <br>\nAnyway, when I can I will investigate it further and I will tell you something, I don't know if it is something common or not</p>",
      "rawMarkdown": "Hi @radek1 ! Glad to see you here.\n\nI was taking a look at the data and I have seen that for example **session 0** in train data lasts from **July 31 (ts → 1659304800025)** to **August 28 (ts → 1661684983707)**\n\nIn my concept of a session, lasting a month seems too much, I don't know what you think. \nAnyway, when I can I will investigate it further and I will tell you something, I don't know if it is something common or not",
      "votes": null
    },
    {
      "id": "2015687",
      "postDate": "11/03/2022 13:12:23",
      "content": "<p>Please take a look at the answer <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486\" target=\"_blank\">here</a> or alternatively at the section that discusses this <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">in my EDA</a>.</p>\n<p>This seemed suspect to me as well but turns out it is just about the session definition used by the organizer 🙂</p>",
      "rawMarkdown": "Please take a look at the answer [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486) or alternatively at the section that discusses this [in my EDA](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset).\n\nThis seemed suspect to me as well but turns out it is just about the session definition used by the organizer 🙂",
      "votes": null
    },
    {
      "id": "2016173",
      "postDate": "11/03/2022 20:28:01",
      "content": "<p>As I wrote in yours \"last 20 aids\": well-done, great contribution to Otto competition.</p>",
      "rawMarkdown": "As I wrote in yours \"last 20 aids\": well-done, great contribution to Otto competition.",
      "votes": null
    },
    {
      "id": "2016245",
      "postDate": "11/03/2022 21:28:28",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a>! 😊 </p>",
      "rawMarkdown": "Thank you, @mpwolke! 😊",
      "votes": null
    },
    {
      "id": "2018373",
      "postDate": "11/05/2022 16:19:33",
      "content": "<p>Great work! upvoted!</p>",
      "rawMarkdown": "Great work! upvoted!",
      "votes": null
    },
    {
      "id": "2018688",
      "postDate": "11/05/2022 22:32:22",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/nyamashina\" target=\"_blank\">@nyamashina</a>, very glad you are finding this useful! 🙂</p>",
      "rawMarkdown": "Thank you, @nyamashina, very glad you are finding this useful! 🙂",
      "votes": null
    },
    {
      "id": "2018761",
      "postDate": "11/06/2022 01:18:27",
      "content": "<p>thanks your notebook, i am a new kaggler, cay i ask you , why json to csv or parquet?</p>",
      "rawMarkdown": "thanks your notebook, i am a new kaggler, cay i ask you , why json to csv or parquet?",
      "votes": null
    },
    {
      "id": "2018854",
      "postDate": "11/06/2022 04:44:45",
      "content": "<p>hey <a href=\"https://www.kaggle.com/iocmeta\" target=\"_blank\">@iocmeta</a>! <code>csv</code> or <code>parquet</code> is easier to read and process with the most common library that most people use, which is <code>pandas</code> 🙂 Generally, <code>csv</code> and <code>parquet</code> files are easier to use for ML work in most circumstances, I believe.</p>",
      "rawMarkdown": "hey @iocmeta! `csv` or `parquet` is easier to read and process with the most common library that most people use, which is `pandas` 🙂 Generally, `csv` and `parquet` files are easier to use for ML work in most circumstances, I believe.",
      "votes": null
    },
    {
      "id": "2023012",
      "postDate": "11/09/2022 13:28:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thank you so much for sharing! </p>\n<p>As you said above to Kaggle does not provide enough RAM (30G in CPU)  to process the data, therefore I won't have luck here to run your process_data.ipynb in my apple M1 8G RAM laptop. </p>\n<p>Although I can't process the entire OTTO dataset with your process_data.ipynb, I should be able to run all of the code for experiment as long as I have the entire OTTO dataset stored in my local disk, right?</p>",
      "rawMarkdown": "Hi @radek1 Thank you so much for sharing! \n\nAs you said above to Kaggle does not provide enough RAM (30G in CPU)  to process the data, therefore I won't have luck here to run your process_data.ipynb in my apple M1 8G RAM laptop. \n\nAlthough I can't process the entire OTTO dataset with your process_data.ipynb, I should be able to run all of the code for experiment as long as I have the entire OTTO dataset stored in my local disk, right?",
      "votes": null
    },
    {
      "id": "2023729",
      "postDate": "11/10/2022 01:41:00",
      "content": "<p>Yeah that is a bit of a problem I am afraid. 8GB is very little, you would have to go to extremely lengths to conserve memory.</p>\n<p>I am not completely sure that it is worth it to figure out how to do stuff with that small amount of RAM. I mean, you still can, but you might be better off just running things on Kaggle. 30GB of RAM plus unlimited runtime for CPU only kernels sounds like a sweet deal 🙂</p>\n<p>Of course, the situation doesn't look that hot when it comes to GPUs, but even 40 hours/week can get you very far I feel. </p>\n<p>Essentially -- I think your time might be better spent figuring out how to gain access to a machine with more RAM (be that Kaggle kernels or cloud compute) then trying to figure out how to work with 8GB of RAM (unless you want to learn something really low level like Rust or C, etc, which is super cool as well but maybe a bit orthogonal to maximizing your learning as a data science practitioner building up their portfolio 🙂) </p>",
      "rawMarkdown": "Yeah that is a bit of a problem I am afraid. 8GB is very little, you would have to go to extremely lengths to conserve memory.\n\nI am not completely sure that it is worth it to figure out how to do stuff with that small amount of RAM. I mean, you still can, but you might be better off just running things on Kaggle. 30GB of RAM plus unlimited runtime for CPU only kernels sounds like a sweet deal 🙂\n\nOf course, the situation doesn't look that hot when it comes to GPUs, but even 40 hours/week can get you very far I feel. \n\nEssentially -- I think your time might be better spent figuring out how to gain access to a machine with more RAM (be that Kaggle kernels or cloud compute) then trying to figure out how to work with 8GB of RAM (unless you want to learn something really low level like Rust or C, etc, which is super cool as well but maybe a bit orthogonal to maximizing your learning as a data science practitioner building up their portfolio 🙂)",
      "votes": null
    },
    {
      "id": "2024279",
      "postDate": "11/10/2022 11:51:44",
      "content": "<p>I tried to run <code>process_data.ipynb</code> to find out what the original json file look like. I love your codes, they are nice and simple to read and learn! </p>\n<p>I also wanted to see how much ram does it take on Kaggle to convert test dataset to df. It seems 400MB jsonl takes nearly 4GB to do the processing. </p>\n<p>However, when I reduced the <code>chunksize</code> to 100, it takes longer but the ram also reduced to 1.6 GB, which is 4 times larger than the test data size. It means 11 GB would require 44 GB ram on Kaggle to process. </p>\n<p>So, I guess this is why the processing can not be done on kaggle on the training set. </p>",
      "rawMarkdown": "I tried to run `process_data.ipynb` to find out what the original json file look like. I love your codes, they are nice and simple to read and learn! \n\nI also wanted to see how much ram does it take on Kaggle to convert test dataset to df. It seems 400MB jsonl takes nearly 4GB to do the processing. \n\nHowever, when I reduced the `chunksize` to 100, it takes longer but the ram also reduced to 1.6 GB, which is 4 times larger than the test data size. It means 11 GB would require 44 GB ram on Kaggle to process. \n\nSo, I guess this is why the processing can not be done on kaggle on the training set.",
      "votes": null
    },
    {
      "id": "2024286",
      "postDate": "11/10/2022 11:54:11",
      "content": "<p>Thank you so much Radek for the advices! Very helpful indeed!</p>\n<p>I only saw this reply just now. I should turn on the notification earlier. </p>",
      "rawMarkdown": "Thank you so much Radek for the advices! Very helpful indeed!\n\nI only saw this reply just now. I should turn on the notification earlier.",
      "votes": null
    },
    {
      "id": "2031088",
      "postDate": "11/15/2022 20:32:32",
      "content": "<p>Great work! Updated!</p>",
      "rawMarkdown": "Great work! Updated!",
      "votes": null
    },
    {
      "id": "2031091",
      "postDate": "11/15/2022 20:33:57",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>, appreciate it! 🙏 </p>",
      "rawMarkdown": "Thanks @leiwong, appreciate it! 🙏",
      "votes": null
    },
    {
      "id": "2031393",
      "postDate": "11/16/2022 03:47:39",
      "content": "<p>Thanks so much for sharing this file Radek.  Glad to see how much time I might expect to process the data on my machine!</p>",
      "rawMarkdown": "Thanks so much for sharing this file Radek.  Glad to see how much time I might expect to process the data on my machine!",
      "votes": null
    },
    {
      "id": "2031397",
      "postDate": "11/16/2022 03:50:11",
      "content": "<p>Very glad you are finding this useful, <a href=\"https://www.kaggle.com/mattrosinski\" target=\"_blank\">@mattrosinski</a>! 🙂</p>",
      "rawMarkdown": "Very glad you are finding this useful, @mattrosinski! 🙂",
      "votes": null
    },
    {
      "id": "2034437",
      "postDate": "11/18/2022 06:12:11",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, since you have updated the otto fully optimized dataset in parquet by reducing the RAM further by dividing <code>ts</code> with 1000, I guess the original process_data.ipynb is not updated.</p>\n<p>The updated <code>process_data.ipynb</code> will add this line into the last cell above right under the line <code>test_df.type = test_df.type.astype(np.uint8)</code>, right? </p>\n<pre><code>test_df.ts = (test_df.ts/).astype(np.int32)\n</code></pre>",
      "rawMarkdown": "Hi @radek1, since you have updated the otto fully optimized dataset in parquet by reducing the RAM further by dividing `ts` with 1000, I guess the original process_data.ipynb is not updated.\n\nThe updated `process_data.ipynb` will add this line into the last cell above right under the line `test_df.type = test_df.type.astype(np.uint8)`, right? \n```python\ntest_df.ts = (test_df.ts/1000).astype(np.int32)\n```",
      "votes": null
    },
    {
      "id": "2065740",
      "postDate": "12/15/2022 03:11:01",
      "content": "<p>Is this a Full data? Are you using full data running this notebook ?</p>",
      "rawMarkdown": "Is this a Full data? Are you using full data running this notebook ?",
      "votes": null
    },
    {
      "id": "2065755",
      "postDate": "12/15/2022 03:39:57",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/luyi564\" target=\"_blank\">@luyi564</a>! Yup, this is the full dataset</p>",
      "rawMarkdown": "Hey @luyi564! Yup, this is the full dataset",
      "votes": null
    },
    {
      "id": "2066696",
      "postDate": "12/16/2022 00:50:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> , thanks for sharing. Do you still use the dataset you shared <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> for offline training &amp; validation?</p>",
      "rawMarkdown": "Hi @radek1 , thanks for sharing. Do you still use the dataset you shared [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991) for offline training & validation?",
      "votes": null
    },
    {
      "id": "2066697",
      "postDate": "12/16/2022 00:54:36",
      "content": "<p>In this competition, the <code>session</code> is actually <code>user</code></p>",
      "rawMarkdown": "In this competition, the `session` is actually `user`",
      "votes": null
    },
    {
      "id": "2066851",
      "postDate": "12/16/2022 05:57:25",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a>! Yes, while working on the competition I was using the dataset that you mentioned.</p>",
      "rawMarkdown": "Hey @wuwenmin! Yes, while working on the competition I was using the dataset that you mentioned.",
      "votes": null
    },
    {
      "id": "2069432",
      "postDate": "12/19/2022 01:28:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, I just <a href=\"https://www.kaggle.com/code/danielliao/recreate-otto-full-optimized-memory-footprint/\" target=\"_blank\">recreated</a> your optimized memory footprint dataset on Kaggle using polars within 30GB. Thank you for sharing and leading the way!</p>",
      "rawMarkdown": "Hi @radek1, I just [recreated](https://www.kaggle.com/code/danielliao/recreate-otto-full-optimized-memory-footprint/) your optimized memory footprint dataset on Kaggle using polars within 30GB. Thank you for sharing and leading the way!",
      "votes": null
    },
    {
      "id": "2090270",
      "postDate": "01/07/2023 07:19:29",
      "content": "<p>Thanks for your processed dataset!<br>\nYou can also use <code>pd.Dataframe.from_records()</code> to convert list of Dicts (<code>session_data.events</code> here) to a df, and then loop it for all sessions and concat them to get final events df.<br>\nAnyway I think it's more readable in your way to write the code.</p>",
      "rawMarkdown": "Thanks for your processed dataset!\nYou can also use `pd.Dataframe.from_records()` to convert list of Dicts (`session_data.events` here) to a df, and then loop it for all sessions and concat them to get final events df.\nAnyway I think it's more readable in your way to write the code.",
      "votes": null
    },
    {
      "id": "2090601",
      "postDate": "01/07/2023 14:01:19",
      "content": "<p>Thank you for this work. It is very useful to the community.</p>",
      "rawMarkdown": "Thank you for this work. It is very useful to the community.",
      "votes": null
    },
    {
      "id": "2090984",
      "postDate": "01/07/2023 22:06:03",
      "content": "<p>Sorry for the delayed response, unfortunately it is not easy to track new comments using the Kaggle interface…</p>\n<p>But yes, you are right <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a>, one can think of the <code>session</code> being a <code>user id</code>. It is important to keep in mind though that there are no overlapping <code>sessions</code> (or <code>users</code>) between test and train, so using the <code>session id</code> as a feature will note be very useful! </p>",
      "rawMarkdown": "Sorry for the delayed response, unfortunately it is not easy to track new comments using the Kaggle interface...\n\nBut yes, you are right @wuwenmin, one can think of the `session` being a `user id`. It is important to keep in mind though that there are no overlapping `sessions` (or `users`) between test and train, so using the `session id` as a feature will note be very useful!",
      "votes": null
    },
    {
      "id": "2090985",
      "postDate": "01/07/2023 22:07:12",
      "content": "<p><a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a>, thank you so much for your comments, much appreciated 🙂 It is a really nice feeling to have your work be appreciated and see that it is helping others.</p>\n<p>Thank you so much for taking the time to share this feedback with me! 🤗</p>",
      "rawMarkdown": "gpreda, thank you so much for your comments, much appreciated 🙂 It is a really nice feeling to have your work be appreciated and see that it is helping others.\n\nThank you so much for taking the time to share this feedback with me! 🤗",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2015604,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/03/2022 12:16:51",
      "content": "<p>This is the script I used for preprocessing as a screenshot</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fd1438821834f8779624042ba02b4d132%2Fpreprocessing_script.png?generation=1667477805785949&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2034437,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "11/18/2022 06:12:11",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, since you have updated the otto fully optimized dataset in parquet by reducing the RAM further by dividing <code>ts</code> with 1000, I guess the original process_data.ipynb is not updated.</p>\n<p>The updated <code>process_data.ipynb</code> will add this line into the last cell above right under the line <code>test_df.type = test_df.type.astype(np.uint8)</code>, right? </p>\n<pre><code>test_df.ts = (test_df.ts/).astype(np.int32)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2090270,
          "author_name": "lrh123",
          "author_url": "",
          "post_date": "01/07/2023 07:19:29",
          "content": "<p>Thanks for your processed dataset!<br>\nYou can also use <code>pd.Dataframe.from_records()</code> to convert list of Dicts (<code>session_data.events</code> here) to a df, and then loop it for all sessions and concat them to get final events df.<br>\nAnyway I think it's more readable in your way to write the code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2015683,
      "author_name": "trasibulo",
      "author_url": "",
      "post_date": "11/03/2022 13:05:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ! Glad to see you here.</p>\n<p>I was taking a look at the data and I have seen that for example <strong>session 0</strong> in train data lasts from <strong>July 31 (ts → 1659304800025)</strong> to <strong>August 28 (ts → 1661684983707)</strong></p>\n<p>In my concept of a session, lasting a month seems too much, I don't know what you think. <br>\nAnyway, when I can I will investigate it further and I will tell you something, I don't know if it is something common or not</p>",
      "votes": null,
      "replies": [
        {
          "id": 2015687,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 13:12:23",
          "content": "<p>Please take a look at the answer <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486\" target=\"_blank\">here</a> or alternatively at the section that discusses this <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">in my EDA</a>.</p>\n<p>This seemed suspect to me as well but turns out it is just about the session definition used by the organizer 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2066697,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/16/2022 00:54:36",
          "content": "<p>In this competition, the <code>session</code> is actually <code>user</code></p>",
          "votes": null,
          "replies": [
            {
              "id": 2090984,
              "author_name": "radek1",
              "author_url": "",
              "post_date": "01/07/2023 22:06:03",
              "content": "<p>Sorry for the delayed response, unfortunately it is not easy to track new comments using the Kaggle interface…</p>\n<p>But yes, you are right <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a>, one can think of the <code>session</code> being a <code>user id</code>. It is important to keep in mind though that there are no overlapping <code>sessions</code> (or <code>users</code>) between test and train, so using the <code>session id</code> as a feature will note be very useful! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2016173,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "11/03/2022 20:28:01",
      "content": "<p>As I wrote in yours \"last 20 aids\": well-done, great contribution to Otto competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2016245,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 21:28:28",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a>! 😊 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2018373,
      "author_name": "nyamashina",
      "author_url": "",
      "post_date": "11/05/2022 16:19:33",
      "content": "<p>Great work! upvoted!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2018688,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/05/2022 22:32:22",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/nyamashina\" target=\"_blank\">@nyamashina</a>, very glad you are finding this useful! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2018761,
      "author_name": "iocmeta",
      "author_url": "",
      "post_date": "11/06/2022 01:18:27",
      "content": "<p>thanks your notebook, i am a new kaggler, cay i ask you , why json to csv or parquet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2018854,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/06/2022 04:44:45",
          "content": "<p>hey <a href=\"https://www.kaggle.com/iocmeta\" target=\"_blank\">@iocmeta</a>! <code>csv</code> or <code>parquet</code> is easier to read and process with the most common library that most people use, which is <code>pandas</code> 🙂 Generally, <code>csv</code> and <code>parquet</code> files are easier to use for ML work in most circumstances, I believe.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2023012,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/09/2022 13:28:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thank you so much for sharing! </p>\n<p>As you said above to Kaggle does not provide enough RAM (30G in CPU)  to process the data, therefore I won't have luck here to run your process_data.ipynb in my apple M1 8G RAM laptop. </p>\n<p>Although I can't process the entire OTTO dataset with your process_data.ipynb, I should be able to run all of the code for experiment as long as I have the entire OTTO dataset stored in my local disk, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2023729,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/10/2022 01:41:00",
          "content": "<p>Yeah that is a bit of a problem I am afraid. 8GB is very little, you would have to go to extremely lengths to conserve memory.</p>\n<p>I am not completely sure that it is worth it to figure out how to do stuff with that small amount of RAM. I mean, you still can, but you might be better off just running things on Kaggle. 30GB of RAM plus unlimited runtime for CPU only kernels sounds like a sweet deal 🙂</p>\n<p>Of course, the situation doesn't look that hot when it comes to GPUs, but even 40 hours/week can get you very far I feel. </p>\n<p>Essentially -- I think your time might be better spent figuring out how to gain access to a machine with more RAM (be that Kaggle kernels or cloud compute) then trying to figure out how to work with 8GB of RAM (unless you want to learn something really low level like Rust or C, etc, which is super cool as well but maybe a bit orthogonal to maximizing your learning as a data science practitioner building up their portfolio 🙂) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2024286,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "11/10/2022 11:54:11",
          "content": "<p>Thank you so much Radek for the advices! Very helpful indeed!</p>\n<p>I only saw this reply just now. I should turn on the notification earlier. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2024279,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/10/2022 11:51:44",
      "content": "<p>I tried to run <code>process_data.ipynb</code> to find out what the original json file look like. I love your codes, they are nice and simple to read and learn! </p>\n<p>I also wanted to see how much ram does it take on Kaggle to convert test dataset to df. It seems 400MB jsonl takes nearly 4GB to do the processing. </p>\n<p>However, when I reduced the <code>chunksize</code> to 100, it takes longer but the ram also reduced to 1.6 GB, which is 4 times larger than the test data size. It means 11 GB would require 44 GB ram on Kaggle to process. </p>\n<p>So, I guess this is why the processing can not be done on kaggle on the training set. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031088,
      "author_name": "leiwong",
      "author_url": "",
      "post_date": "11/15/2022 20:32:32",
      "content": "<p>Great work! Updated!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031091,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/15/2022 20:33:57",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>, appreciate it! 🙏 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031393,
      "author_name": "mattrosinski",
      "author_url": "",
      "post_date": "11/16/2022 03:47:39",
      "content": "<p>Thanks so much for sharing this file Radek.  Glad to see how much time I might expect to process the data on my machine!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031397,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/16/2022 03:50:11",
          "content": "<p>Very glad you are finding this useful, <a href=\"https://www.kaggle.com/mattrosinski\" target=\"_blank\">@mattrosinski</a>! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2065740,
      "author_name": "luyi564",
      "author_url": "",
      "post_date": "12/15/2022 03:11:01",
      "content": "<p>Is this a Full data? Are you using full data running this notebook ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2065755,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/15/2022 03:39:57",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/luyi564\" target=\"_blank\">@luyi564</a>! Yup, this is the full dataset</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2066696,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "12/16/2022 00:50:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> , thanks for sharing. Do you still use the dataset you shared <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> for offline training &amp; validation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2066851,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/16/2022 05:57:25",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a>! Yes, while working on the competition I was using the dataset that you mentioned.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2069432,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "12/19/2022 01:28:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, I just <a href=\"https://www.kaggle.com/code/danielliao/recreate-otto-full-optimized-memory-footprint/\" target=\"_blank\">recreated</a> your optimized memory footprint dataset on Kaggle using polars within 30GB. Thank you for sharing and leading the way!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2090601,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "01/07/2023 14:01:19",
      "content": "<p>Thank you for this work. It is very useful to the community.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2090985,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "01/07/2023 22:07:12",
          "content": "<p><a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a>, thank you so much for your comments, much appreciated 🙂 It is a really nice feeling to have your work be appreciated and see that it is helping others.</p>\n<p>Thank you so much for taking the time to share this feedback with me! 🤗</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2015554": "Hey friends!\n\nI converted the dataset to live in `csv` & `parquet` files. You can find the dataset [here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).\n\nI have also created [an explainer notebook](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-file). In it, I run you through how to use this data.\n\nI am also attaching the code I used to process this data as a notebook file to this post. I couldn't process the data on Kaggle due to RAM limitations so I ran it on my local machine and uploaded it to Kaggle.\n\nHope this can speed you along in your work 🙂 Happy kaggling! 🥳\n\n\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2015604": "This is the script I used for preprocessing as a screenshot\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fd1438821834f8779624042ba02b4d132%2Fpreprocessing_script.png?generation=1667477805785949&alt=media)",
    "2015683": "Hi @radek1 ! Glad to see you here.\n\nI was taking a look at the data and I have seen that for example **session 0** in train data lasts from **July 31 (ts → 1659304800025)** to **August 28 (ts → 1661684983707)**\n\nIn my concept of a session, lasting a month seems too much, I don't know what you think. \nAnyway, when I can I will investigate it further and I will tell you something, I don't know if it is something common or not",
    "2015687": "Please take a look at the answer [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486) or alternatively at the section that discusses this [in my EDA](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset).\n\nThis seemed suspect to me as well but turns out it is just about the session definition used by the organizer 🙂",
    "2016173": "As I wrote in yours \"last 20 aids\": well-done, great contribution to Otto competition.",
    "2016245": "Thank you, @mpwolke! 😊",
    "2018373": "Great work! upvoted!",
    "2018688": "Thank you, @nyamashina, very glad you are finding this useful! 🙂",
    "2018761": "thanks your notebook, i am a new kaggler, cay i ask you , why json to csv or parquet?",
    "2018854": "hey @iocmeta! `csv` or `parquet` is easier to read and process with the most common library that most people use, which is `pandas` 🙂 Generally, `csv` and `parquet` files are easier to use for ML work in most circumstances, I believe.",
    "2023012": "Hi @radek1 Thank you so much for sharing! \n\nAs you said above to Kaggle does not provide enough RAM (30G in CPU)  to process the data, therefore I won't have luck here to run your process_data.ipynb in my apple M1 8G RAM laptop. \n\nAlthough I can't process the entire OTTO dataset with your process_data.ipynb, I should be able to run all of the code for experiment as long as I have the entire OTTO dataset stored in my local disk, right?",
    "2023729": "Yeah that is a bit of a problem I am afraid. 8GB is very little, you would have to go to extremely lengths to conserve memory.\n\nI am not completely sure that it is worth it to figure out how to do stuff with that small amount of RAM. I mean, you still can, but you might be better off just running things on Kaggle. 30GB of RAM plus unlimited runtime for CPU only kernels sounds like a sweet deal 🙂\n\nOf course, the situation doesn't look that hot when it comes to GPUs, but even 40 hours/week can get you very far I feel. \n\nEssentially -- I think your time might be better spent figuring out how to gain access to a machine with more RAM (be that Kaggle kernels or cloud compute) then trying to figure out how to work with 8GB of RAM (unless you want to learn something really low level like Rust or C, etc, which is super cool as well but maybe a bit orthogonal to maximizing your learning as a data science practitioner building up their portfolio 🙂)",
    "2024279": "I tried to run `process_data.ipynb` to find out what the original json file look like. I love your codes, they are nice and simple to read and learn! \n\nI also wanted to see how much ram does it take on Kaggle to convert test dataset to df. It seems 400MB jsonl takes nearly 4GB to do the processing. \n\nHowever, when I reduced the `chunksize` to 100, it takes longer but the ram also reduced to 1.6 GB, which is 4 times larger than the test data size. It means 11 GB would require 44 GB ram on Kaggle to process. \n\nSo, I guess this is why the processing can not be done on kaggle on the training set.",
    "2024286": "Thank you so much Radek for the advices! Very helpful indeed!\n\nI only saw this reply just now. I should turn on the notification earlier.",
    "2031088": "Great work! Updated!",
    "2031091": "Thanks @leiwong, appreciate it! 🙏",
    "2031393": "Thanks so much for sharing this file Radek.  Glad to see how much time I might expect to process the data on my machine!",
    "2031397": "Very glad you are finding this useful, @mattrosinski! 🙂",
    "2034437": "Hi @radek1, since you have updated the otto fully optimized dataset in parquet by reducing the RAM further by dividing `ts` with 1000, I guess the original process_data.ipynb is not updated.\n\nThe updated `process_data.ipynb` will add this line into the last cell above right under the line `test_df.type = test_df.type.astype(np.uint8)`, right? \n```python\ntest_df.ts = (test_df.ts/1000).astype(np.int32)\n```",
    "2065740": "Is this a Full data? Are you using full data running this notebook ?",
    "2065755": "Hey @luyi564! Yup, this is the full dataset",
    "2066696": "Hi @radek1 , thanks for sharing. Do you still use the dataset you shared [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991) for offline training & validation?",
    "2066697": "In this competition, the `session` is actually `user`",
    "2066851": "Hey @wuwenmin! Yes, while working on the competition I was using the dataset that you mentioned.",
    "2069432": "Hi @radek1, I just [recreated](https://www.kaggle.com/code/danielliao/recreate-otto-full-optimized-memory-footprint/) your optimized memory footprint dataset on Kaggle using polars within 30GB. Thank you for sharing and leading the way!",
    "2090270": "Thanks for your processed dataset!\nYou can also use `pd.Dataframe.from_records()` to convert list of Dicts (`session_data.events` here) to a df, and then loop it for all sessions and concat them to get final events df.\nAnyway I think it's more readable in your way to write the code.",
    "2090601": "Thank you for this work. It is very useful to the community.",
    "2090984": "Sorry for the delayed response, unfortunately it is not easy to track new comments using the Kaggle interface...\n\nBut yes, you are right @wuwenmin, one can think of the `session` being a `user id`. It is important to keep in mind though that there are no overlapping `sessions` (or `users`) between test and train, so using the `session id` as a feature will note be very useful!",
    "2090985": "gpreda, thank you so much for your comments, much appreciated 🙂 It is a really nice feeling to have your work be appreciated and see that it is helping others.\n\nThank you so much for taking the time to share this feedback with me! 🤗"
  },
  "source": "meta"
}