{
  "id": 308423,
  "title": "Is PySpark better tool for this dataset size ?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308423",
  "author_name": "",
  "post_date": "2022-02-18T15:07:47.785143200Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This dataset is not small volume. And I am thinking to use PySpark for preprocessing.<br>\nOf course, we can also handle this data volume by Pandas.<br>\nI want to know which tool is best for this dataset.   </p>",
  "messages": [
    {
      "id": "1696066",
      "postDate": "02/18/2022 15:07:47",
      "content": "<p>This dataset is not small volume. And I am thinking to use PySpark for preprocessing.<br>\nOf course, we can also handle this data volume by Pandas.<br>\nI want to know which tool is best for this dataset.   </p>",
      "rawMarkdown": "This dataset is not small volume. And I am thinking to use PySpark for preprocessing.\nOf course, we can also handle this data volume by Pandas.\nI want to know which tool is best for this dataset.",
      "votes": null
    },
    {
      "id": "1696078",
      "postDate": "02/18/2022 15:14:43",
      "content": "<p><a href=\"https://www.kaggle.com/sifury\" target=\"_blank\">@sifury</a> you should also checkout rapids.ai and NVIDIA Merlin ecosystem </p>",
      "rawMarkdown": "sifury you should also checkout rapids.ai and NVIDIA Merlin ecosystem",
      "votes": null
    },
    {
      "id": "1696199",
      "postDate": "02/18/2022 16:47:39",
      "content": "<p>In my experience, Merlin is still not good enough for real applications (environment, data preprocessing &amp; scheme, memory issues both with preprocessing and training&amp;validation loop (tr4rec), troubles with inference) + you need a lot of GPU memory.</p>",
      "rawMarkdown": "In my experience, Merlin is still not good enough for real applications (environment, data preprocessing & scheme, memory issues both with preprocessing and training&validation loop (tr4rec), troubles with inference) + you need a lot of GPU memory.",
      "votes": null
    },
    {
      "id": "1696242",
      "postDate": "02/18/2022 17:26:58",
      "content": "<p>I was thinking the same, but I found few ways to deal with the dataset, at least for the tabular part : </p>\n<ul>\n<li><p>Use another format than csv, like Parquet, for instance transactions.csv is around 3GB but in parquet format it's only 800mb which is 200% less size and loads 50% faster with pandas <code>read_parquet</code> , you can experiment with other formats.</p></li>\n<li><p>Try to use pure pythons dicts to map some stuff that you need a lot, this will give a lot of performance, this is basically how large databases operates.</p></li>\n<li><p>Try to vectorize your pandas operations (avoid iterating on the dataframe for instance, use agg and apply if you cant' fully vectorize).</p></li>\n<li><p>Use NumPy whenever you can for heavy mathematical operations.</p></li>\n</ul>",
      "rawMarkdown": "I was thinking the same, but I found few ways to deal with the dataset, at least for the tabular part : \n\n- Use another format than csv, like Parquet, for instance transactions.csv is around 3GB but in parquet format it's only 800mb which is 200% less size and loads 50% faster with pandas `read_parquet` , you can experiment with other formats.\n\n- Try to use pure pythons dicts to map some stuff that you need a lot, this will give a lot of performance, this is basically how large databases operates.\n\n- Try to vectorize your pandas operations (avoid iterating on the dataframe for instance, use agg and apply if you cant' fully vectorize).\n\n- Use NumPy whenever you can for heavy mathematical operations.",
      "votes": null
    },
    {
      "id": "1696356",
      "postDate": "02/18/2022 19:10:54",
      "content": "<p>I wasn't aware of this thanks for sharing! I believe the standard requirement is &gt;=V100 which might be what most people would utilize? </p>\n<p>Could you please elaborate on what issues you've run into? I was under the impression it might be a great spin/experiment for this competition</p>",
      "rawMarkdown": "I wasn't aware of this thanks for sharing! I believe the standard requirement is >=V100 which might be what most people would utilize? \n\nCould you please elaborate on what issues you've run into? I was under the impression it might be a great spin/experiment for this competition",
      "votes": null
    },
    {
      "id": "1698093",
      "postDate": "02/20/2022 06:12:29",
      "content": "<p>Chris has explained a <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">neat memory trick here</a></p>",
      "rawMarkdown": "Chris has explained a [neat memory trick here](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635)",
      "votes": null
    },
    {
      "id": "1698300",
      "postDate": "02/20/2022 09:19:43",
      "content": "<p>Thanks for excellent tips. I just found parquet input format.<br>\n<a href=\"https://www.kaggle.com/sytuannguyen/hm2022-low-memory-fast-loading\" target=\"_blank\">https://www.kaggle.com/sytuannguyen/hm2022-low-memory-fast-loading</a></p>",
      "rawMarkdown": "Thanks for excellent tips. I just found parquet input format.\nhttps://www.kaggle.com/sytuannguyen/hm2022-low-memory-fast-loading",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1696078,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/18/2022 15:14:43",
      "content": "<p><a href=\"https://www.kaggle.com/sifury\" target=\"_blank\">@sifury</a> you should also checkout rapids.ai and NVIDIA Merlin ecosystem </p>",
      "votes": null,
      "replies": [
        {
          "id": 1696199,
          "author_name": "simakov",
          "author_url": "",
          "post_date": "02/18/2022 16:47:39",
          "content": "<p>In my experience, Merlin is still not good enough for real applications (environment, data preprocessing &amp; scheme, memory issues both with preprocessing and training&amp;validation loop (tr4rec), troubles with inference) + you need a lot of GPU memory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696356,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/18/2022 19:10:54",
          "content": "<p>I wasn't aware of this thanks for sharing! I believe the standard requirement is &gt;=V100 which might be what most people would utilize? </p>\n<p>Could you please elaborate on what issues you've run into? I was under the impression it might be a great spin/experiment for this competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1696242,
      "author_name": "souamesannis",
      "author_url": "",
      "post_date": "02/18/2022 17:26:58",
      "content": "<p>I was thinking the same, but I found few ways to deal with the dataset, at least for the tabular part : </p>\n<ul>\n<li><p>Use another format than csv, like Parquet, for instance transactions.csv is around 3GB but in parquet format it's only 800mb which is 200% less size and loads 50% faster with pandas <code>read_parquet</code> , you can experiment with other formats.</p></li>\n<li><p>Try to use pure pythons dicts to map some stuff that you need a lot, this will give a lot of performance, this is basically how large databases operates.</p></li>\n<li><p>Try to vectorize your pandas operations (avoid iterating on the dataframe for instance, use agg and apply if you cant' fully vectorize).</p></li>\n<li><p>Use NumPy whenever you can for heavy mathematical operations.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1698093,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/20/2022 06:12:29",
          "content": "<p>Chris has explained a <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">neat memory trick here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698300,
          "author_name": "sifury",
          "author_url": "",
          "post_date": "02/20/2022 09:19:43",
          "content": "<p>Thanks for excellent tips. I just found parquet input format.<br>\n<a href=\"https://www.kaggle.com/sytuannguyen/hm2022-low-memory-fast-loading\" target=\"_blank\">https://www.kaggle.com/sytuannguyen/hm2022-low-memory-fast-loading</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1696066": "This dataset is not small volume. And I am thinking to use PySpark for preprocessing.\nOf course, we can also handle this data volume by Pandas.\nI want to know which tool is best for this dataset.",
    "1696078": "sifury you should also checkout rapids.ai and NVIDIA Merlin ecosystem",
    "1696199": "In my experience, Merlin is still not good enough for real applications (environment, data preprocessing & scheme, memory issues both with preprocessing and training&validation loop (tr4rec), troubles with inference) + you need a lot of GPU memory.",
    "1696242": "I was thinking the same, but I found few ways to deal with the dataset, at least for the tabular part : \n\n- Use another format than csv, like Parquet, for instance transactions.csv is around 3GB but in parquet format it's only 800mb which is 200% less size and loads 50% faster with pandas `read_parquet` , you can experiment with other formats.\n\n- Try to use pure pythons dicts to map some stuff that you need a lot, this will give a lot of performance, this is basically how large databases operates.\n\n- Try to vectorize your pandas operations (avoid iterating on the dataframe for instance, use agg and apply if you cant' fully vectorize).\n\n- Use NumPy whenever you can for heavy mathematical operations.",
    "1696356": "I wasn't aware of this thanks for sharing! I believe the standard requirement is >=V100 which might be what most people would utilize? \n\nCould you please elaborate on what issues you've run into? I was under the impression it might be a great spin/experiment for this competition",
    "1698093": "Chris has explained a [neat memory trick here](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635)",
    "1698300": "Thanks for excellent tips. I just found parquet input format.\nhttps://www.kaggle.com/sytuannguyen/hm2022-low-memory-fast-loading"
  },
  "source": "meta"
}