{
  "id": 475485,
  "title": "Best Practices of Handling BigData-Scale Datasets in Code Competitions",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475485",
  "author_name": "",
  "post_date": "2024-02-08T16:27:38.552181700Z",
  "votes": 106,
  "comment_count": 19,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>The contest of “Home Credit - Credit Risk Model Stability” imposes us to operate under the significant resource constraints.</p>\n<ul>\n<li>Handling the large (really BigData-scale) dataset</li>\n<li>Running it as a code competition on Kaggle notebooks</li>\n</ul>\n<p>The above-mentioned circumstances force us to invent solutions that are</p>\n<ul>\n<li>Smart in terms of utilizing RAM, CPU/GPU/TPU time, and disk storage. </li>\n<li>Fast in terms of the execution time (basic Kaggle kernels have the limit of a session to run up to 9 h, with visually reported execution time not exceeding 6 h).</li>\n</ul>\n<p>Even if you decide to convert your Kaggle notebook into a Google Collab Pro+ one (it is possible but it comes at a price: <a href=\"https://colab.research.google.com/signup)\" target=\"_blank\">https://colab.research.google.com/signup)</a>, it would not resolve all of the resource constraint-driven problems for your codebase.</p>\n<p>Therefore, the only viable option is to get your head around the best practices of the effective tackling BigData-scale problems in the online resource-limited kernels provided by Kaggle. Some of the best practices are universal, and others are specific to the type of operations you try to accomplish (like EDA-specific or ML Engineering-specific tricks). We are going to review them in the sections below.</p>\n<h1>Universal tricks</h1>\n<ul>\n<li>Load your data in fast, time-efficient manner (yes, reading BigData with Pandas directly is a pain as well as the drain of the execution time of your session, even if you utilize a batch reading approach); the good example of a fast data loading strategy with <em>polar</em> is provided in <a href=\"https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive\" target=\"_blank\">https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive</a> ; obviously, this is not the only strategy possible to implement – you should be smart to look at the appropriate options (<em>pyarrrow</em> to be one of the things to look at, besides <em>polar</em>; btw, the nice and fast <em>pyarrow</em>-based data loading strategy demonstrated by <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> in <a href=\"https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet\" target=\"_blank\">https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet</a>)</li>\n<li>Optimize the in-memory size of Pandas dataframe (some of the techniques of this sort are shown in <a href=\"https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive\" target=\"_blank\">https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive</a>, for instance)</li>\n<li>Force regular garbage collection in the  certain points of your notebook, using the capabilities of <em>gc</em> package</li>\n<li>Store the preprocessed training data as a separate Kaggle dataset – it will save the execution of your notebooks doing EDA or/and ML experiments down the road</li>\n</ul>\n<p>The good options for the time-efficient loading could be</p>\n<ul>\n<li><em>polars</em></li>\n<li><em>dask</em></li>\n<li><em>datatable</em></li>\n<li>novel pandas 2 with pyarrow and schema definition (the good trick of this sort can be reused from the demo in <a href=\"https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet\" target=\"_blank\">this notebook</a> implemented for another contest with BigData-scale data processing requirements)</li>\n<li><em>rapids</em></li>\n<li>etc.</li>\n</ul>\n<p><strong>Note 1:</strong> you can refer to the excellent post <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475389\" target=\"_blank\">Faster ways to load competition's data</a> by <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> to see the benchmarking of some of the above-mentioned tools as applied to the data in the dataset for this competition</p>\n<p><strong>Note 2:</strong> Thanks to <a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> for mentioning <em><a href=\"https://rapids.ai/\" target=\"_blank\">rapids</a></em>. Brief scanning through its documentation as well as looking through the Kaggle-tailored <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">tutorial</a> displayed a lot of potential in it.</p>\n<h1>EDA-specific tricks</h1>\n<p>Plotting the datasets that are bigger them a RAM memory available to your solution will crash your conventional workflows (like trying to feed a pandas dataframe to a graph lib like plotly or seaborn and then visualize something). </p>\n<p>Therefore, more elaborate approaches should be used to achieve your goal. One of the popular patterns of doing it is described below</p>\n<ul>\n<li>Launch a memory-efficient lightweight database engine (like DuckDB etc.)</li>\n<li>Mount or import your BigData-scale files into it as tables</li>\n<li>Pre-calculate the data needed for your charts (histograms, boxplots, bar charts etc.) with SQL queries</li>\n<li>Visualize your pre-calculated data with a graph library of your choice (with minimal client-side calculations)</li>\n</ul>\n<p>The good case of implementing such a strategy is provided in <a href=\"https://www.kaggle.com/code/gvyshnya/plotting-big-data-sleeping-series\" target=\"_blank\">Plotting Big Data: Sleeping Series</a> where DuckDB and Plotly were used together to cook the noodles (it is not for this very dataset but the ideas from there could be easily transferred to this data domain).</p>\n<p>SQLLite can be also leveraged in such a situation (see <a href=\"https://plotly.com/python/v3/ipython-notebooks/big-data-analytics-with-pandas-and-sqlite/\" target=\"_blank\">here</a> as an example).</p>\n<h1>Specific ML modelling tricks</h1>\n<p>Exercising every general best practice described in the respective section above is essential for the time-efficient ML pipeline built on Kaggle kernels. However, there is one more trick that can help you. </p>\n<p>You can</p>\n<ul>\n<li>Train your model out of the Kaggle premise</li>\n<li>Save (serialize) it as a binary file (either pickle or a native TensorFlow model storage format could be used, depending on what a model you train) – in such a way, you get a pre-trained model in your disposal</li>\n<li>Upload the pre-trained model to a new Kaggle dataset</li>\n<li>Import such a dataset into your notebook intended to submit predictions to load the pre-trained model from the binary file</li>\n</ul>\n<p>Other thing to consider is to unitilize the batch prediction strategy where your model predicts against the subsamples of the test set rather then on the entire test set at once.</p>\n<h1>Shall Pandas Be Abandoned?</h1>\n<p>Even if you become versatile at using BigData-scale platforms like ones mentioned above, you will stilll have to keep <em>pandas</em> in your toolbox. Tasks like clustering or visualizing your data would force a conversion of preprocessed data to pandas first. The matter  is, not every special-purpose lib (graphs, clustering, and even ML) provides support for something like polars or dask out of the box.</p>\n<p>At the same time, fast and memory-efficient data preprocessing will make such a niched pandas utilization less painful to your solutions in terms of its memory and CPU footprints.</p>\n<h1>Final notes</h1>\n<p>If we were not limited to Kaggle kernels in this competition, I would go further to share one more recommendation. It is about taking advantage of <strong>lazy evaluation</strong>. Lazy evaluation refers to the strategy which delays evaluation of an expression until the value is actually needed. Lazy evaluation is an important concept (used especially in functional programming), and if you want to read more about its different usages in Python you can start <a href=\"https://towardsdatascience.com/what-is-lazy-evaluation-in-python-9efb1d3bfed0\" target=\"_blank\">here</a>.</p>\n<p>Many industrial platforms (like Spark/PySpark, Ray etc.) facilitate it for you out of the box. For example, I stick to Spark/PySpark as a go-to choice to build BigData-scale solutions in the industrial projects. </p>\n<p>However, launching the infrastructure needed for the platforms like Spark/PySpark requires some resources (RAM, CPU) to be consumed, and it could make it impossible to use on Kaggle kernels with the datasets of a size compared with the data for this contest.</p>",
  "messages": [
    {
      "id": "2643121",
      "postDate": "02/08/2024 16:27:38",
      "content": "<h1>Introduction</h1>\n<p>The contest of “Home Credit - Credit Risk Model Stability” imposes us to operate under the significant resource constraints.</p>\n<ul>\n<li>Handling the large (really BigData-scale) dataset</li>\n<li>Running it as a code competition on Kaggle notebooks</li>\n</ul>\n<p>The above-mentioned circumstances force us to invent solutions that are</p>\n<ul>\n<li>Smart in terms of utilizing RAM, CPU/GPU/TPU time, and disk storage. </li>\n<li>Fast in terms of the execution time (basic Kaggle kernels have the limit of a session to run up to 9 h, with visually reported execution time not exceeding 6 h).</li>\n</ul>\n<p>Even if you decide to convert your Kaggle notebook into a Google Collab Pro+ one (it is possible but it comes at a price: <a href=\"https://colab.research.google.com/signup)\" target=\"_blank\">https://colab.research.google.com/signup)</a>, it would not resolve all of the resource constraint-driven problems for your codebase.</p>\n<p>Therefore, the only viable option is to get your head around the best practices of the effective tackling BigData-scale problems in the online resource-limited kernels provided by Kaggle. Some of the best practices are universal, and others are specific to the type of operations you try to accomplish (like EDA-specific or ML Engineering-specific tricks). We are going to review them in the sections below.</p>\n<h1>Universal tricks</h1>\n<ul>\n<li>Load your data in fast, time-efficient manner (yes, reading BigData with Pandas directly is a pain as well as the drain of the execution time of your session, even if you utilize a batch reading approach); the good example of a fast data loading strategy with <em>polar</em> is provided in <a href=\"https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive\" target=\"_blank\">https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive</a> ; obviously, this is not the only strategy possible to implement – you should be smart to look at the appropriate options (<em>pyarrrow</em> to be one of the things to look at, besides <em>polar</em>; btw, the nice and fast <em>pyarrow</em>-based data loading strategy demonstrated by <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> in <a href=\"https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet\" target=\"_blank\">https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet</a>)</li>\n<li>Optimize the in-memory size of Pandas dataframe (some of the techniques of this sort are shown in <a href=\"https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive\" target=\"_blank\">https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive</a>, for instance)</li>\n<li>Force regular garbage collection in the  certain points of your notebook, using the capabilities of <em>gc</em> package</li>\n<li>Store the preprocessed training data as a separate Kaggle dataset – it will save the execution of your notebooks doing EDA or/and ML experiments down the road</li>\n</ul>\n<p>The good options for the time-efficient loading could be</p>\n<ul>\n<li><em>polars</em></li>\n<li><em>dask</em></li>\n<li><em>datatable</em></li>\n<li>novel pandas 2 with pyarrow and schema definition (the good trick of this sort can be reused from the demo in <a href=\"https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet\" target=\"_blank\">this notebook</a> implemented for another contest with BigData-scale data processing requirements)</li>\n<li><em>rapids</em></li>\n<li>etc.</li>\n</ul>\n<p><strong>Note 1:</strong> you can refer to the excellent post <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475389\" target=\"_blank\">Faster ways to load competition's data</a> by <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> to see the benchmarking of some of the above-mentioned tools as applied to the data in the dataset for this competition</p>\n<p><strong>Note 2:</strong> Thanks to <a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> for mentioning <em><a href=\"https://rapids.ai/\" target=\"_blank\">rapids</a></em>. Brief scanning through its documentation as well as looking through the Kaggle-tailored <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">tutorial</a> displayed a lot of potential in it.</p>\n<h1>EDA-specific tricks</h1>\n<p>Plotting the datasets that are bigger them a RAM memory available to your solution will crash your conventional workflows (like trying to feed a pandas dataframe to a graph lib like plotly or seaborn and then visualize something). </p>\n<p>Therefore, more elaborate approaches should be used to achieve your goal. One of the popular patterns of doing it is described below</p>\n<ul>\n<li>Launch a memory-efficient lightweight database engine (like DuckDB etc.)</li>\n<li>Mount or import your BigData-scale files into it as tables</li>\n<li>Pre-calculate the data needed for your charts (histograms, boxplots, bar charts etc.) with SQL queries</li>\n<li>Visualize your pre-calculated data with a graph library of your choice (with minimal client-side calculations)</li>\n</ul>\n<p>The good case of implementing such a strategy is provided in <a href=\"https://www.kaggle.com/code/gvyshnya/plotting-big-data-sleeping-series\" target=\"_blank\">Plotting Big Data: Sleeping Series</a> where DuckDB and Plotly were used together to cook the noodles (it is not for this very dataset but the ideas from there could be easily transferred to this data domain).</p>\n<p>SQLLite can be also leveraged in such a situation (see <a href=\"https://plotly.com/python/v3/ipython-notebooks/big-data-analytics-with-pandas-and-sqlite/\" target=\"_blank\">here</a> as an example).</p>\n<h1>Specific ML modelling tricks</h1>\n<p>Exercising every general best practice described in the respective section above is essential for the time-efficient ML pipeline built on Kaggle kernels. However, there is one more trick that can help you. </p>\n<p>You can</p>\n<ul>\n<li>Train your model out of the Kaggle premise</li>\n<li>Save (serialize) it as a binary file (either pickle or a native TensorFlow model storage format could be used, depending on what a model you train) – in such a way, you get a pre-trained model in your disposal</li>\n<li>Upload the pre-trained model to a new Kaggle dataset</li>\n<li>Import such a dataset into your notebook intended to submit predictions to load the pre-trained model from the binary file</li>\n</ul>\n<p>Other thing to consider is to unitilize the batch prediction strategy where your model predicts against the subsamples of the test set rather then on the entire test set at once.</p>\n<h1>Shall Pandas Be Abandoned?</h1>\n<p>Even if you become versatile at using BigData-scale platforms like ones mentioned above, you will stilll have to keep <em>pandas</em> in your toolbox. Tasks like clustering or visualizing your data would force a conversion of preprocessed data to pandas first. The matter  is, not every special-purpose lib (graphs, clustering, and even ML) provides support for something like polars or dask out of the box.</p>\n<p>At the same time, fast and memory-efficient data preprocessing will make such a niched pandas utilization less painful to your solutions in terms of its memory and CPU footprints.</p>\n<h1>Final notes</h1>\n<p>If we were not limited to Kaggle kernels in this competition, I would go further to share one more recommendation. It is about taking advantage of <strong>lazy evaluation</strong>. Lazy evaluation refers to the strategy which delays evaluation of an expression until the value is actually needed. Lazy evaluation is an important concept (used especially in functional programming), and if you want to read more about its different usages in Python you can start <a href=\"https://towardsdatascience.com/what-is-lazy-evaluation-in-python-9efb1d3bfed0\" target=\"_blank\">here</a>.</p>\n<p>Many industrial platforms (like Spark/PySpark, Ray etc.) facilitate it for you out of the box. For example, I stick to Spark/PySpark as a go-to choice to build BigData-scale solutions in the industrial projects. </p>\n<p>However, launching the infrastructure needed for the platforms like Spark/PySpark requires some resources (RAM, CPU) to be consumed, and it could make it impossible to use on Kaggle kernels with the datasets of a size compared with the data for this contest.</p>",
      "rawMarkdown": "# Introduction\n\nThe contest of “Home Credit - Credit Risk Model Stability” imposes us to operate under the significant resource constraints.\n- Handling the large (really BigData-scale) dataset\n- Running it as a code competition on Kaggle notebooks\n\nThe above-mentioned circumstances force us to invent solutions that are\n- Smart in terms of utilizing RAM, CPU/GPU/TPU time, and disk storage. \n- Fast in terms of the execution time (basic Kaggle kernels have the limit of a session to run up to 9 h, with visually reported execution time not exceeding 6 h).\n\nEven if you decide to convert your Kaggle notebook into a Google Collab Pro+ one (it is possible but it comes at a price: https://colab.research.google.com/signup), it would not resolve all of the resource constraint-driven problems for your codebase.\n\nTherefore, the only viable option is to get your head around the best practices of the effective tackling BigData-scale problems in the online resource-limited kernels provided by Kaggle. Some of the best practices are universal, and others are specific to the type of operations you try to accomplish (like EDA-specific or ML Engineering-specific tricks). We are going to review them in the sections below.\n\n# Universal tricks\n\n- Load your data in fast, time-efficient manner (yes, reading BigData with Pandas directly is a pain as well as the drain of the execution time of your session, even if you utilize a batch reading approach); the good example of a fast data loading strategy with *polar* is provided in https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive ; obviously, this is not the only strategy possible to implement – you should be smart to look at the appropriate options (*pyarrrow* to be one of the things to look at, besides *polar*; btw, the nice and fast *pyarrow*-based data loading strategy demonstrated by @tolgadincer in https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet)\n- Optimize the in-memory size of Pandas dataframe (some of the techniques of this sort are shown in https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive, for instance)\n- Force regular garbage collection in the  certain points of your notebook, using the capabilities of *gc* package\n- Store the preprocessed training data as a separate Kaggle dataset – it will save the execution of your notebooks doing EDA or/and ML experiments down the road\n\nThe good options for the time-efficient loading could be\n- *polars*\n- *dask*\n- *datatable*\n- novel pandas 2 with pyarrow and schema definition (the good trick of this sort can be reused from the demo in [this notebook](https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet) implemented for another contest with BigData-scale data processing requirements)\n- *rapids*\n- etc.\n\n**Note 1:** you can refer to the excellent post [Faster ways to load competition's data](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475389) by @kononenko to see the benchmarking of some of the above-mentioned tools as applied to the data in the dataset for this competition\n\n**Note 2:** Thanks to @thomasmeiner for mentioning *[rapids](https://rapids.ai/)*. Brief scanning through its documentation as well as looking through the Kaggle-tailored [tutorial](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets) displayed a lot of potential in it.\n\n# EDA-specific tricks\n\nPlotting the datasets that are bigger them a RAM memory available to your solution will crash your conventional workflows (like trying to feed a pandas dataframe to a graph lib like plotly or seaborn and then visualize something). \n\nTherefore, more elaborate approaches should be used to achieve your goal. One of the popular patterns of doing it is described below\n- Launch a memory-efficient lightweight database engine (like DuckDB etc.)\n- Mount or import your BigData-scale files into it as tables\n- Pre-calculate the data needed for your charts (histograms, boxplots, bar charts etc.) with SQL queries\n- Visualize your pre-calculated data with a graph library of your choice (with minimal client-side calculations)\n\nThe good case of implementing such a strategy is provided in [Plotting Big Data: Sleeping Series](https://www.kaggle.com/code/gvyshnya/plotting-big-data-sleeping-series) where DuckDB and Plotly were used together to cook the noodles (it is not for this very dataset but the ideas from there could be easily transferred to this data domain).\n\nSQLLite can be also leveraged in such a situation (see [here](https://plotly.com/python/v3/ipython-notebooks/big-data-analytics-with-pandas-and-sqlite/) as an example).\n\n# Specific ML modelling tricks\n\nExercising every general best practice described in the respective section above is essential for the time-efficient ML pipeline built on Kaggle kernels. However, there is one more trick that can help you. \n\nYou can\n- Train your model out of the Kaggle premise\n- Save (serialize) it as a binary file (either pickle or a native TensorFlow model storage format could be used, depending on what a model you train) – in such a way, you get a pre-trained model in your disposal\n- Upload the pre-trained model to a new Kaggle dataset\n- Import such a dataset into your notebook intended to submit predictions to load the pre-trained model from the binary file\n\nOther thing to consider is to unitilize the batch prediction strategy where your model predicts against the subsamples of the test set rather then on the entire test set at once.\n\n# Shall Pandas Be Abandoned?\n\nEven if you become versatile at using BigData-scale platforms like ones mentioned above, you will stilll have to keep *pandas* in your toolbox. Tasks like clustering or visualizing your data would force a conversion of preprocessed data to pandas first. The matter  is, not every special-purpose lib (graphs, clustering, and even ML) provides support for something like polars or dask out of the box.\n\nAt the same time, fast and memory-efficient data preprocessing will make such a niched pandas utilization less painful to your solutions in terms of its memory and CPU footprints.\n\n# Final notes\n\nIf we were not limited to Kaggle kernels in this competition, I would go further to share one more recommendation. It is about taking advantage of **lazy evaluation**. Lazy evaluation refers to the strategy which delays evaluation of an expression until the value is actually needed. Lazy evaluation is an important concept (used especially in functional programming), and if you want to read more about its different usages in Python you can start [here](https://towardsdatascience.com/what-is-lazy-evaluation-in-python-9efb1d3bfed0).\n\nMany industrial platforms (like Spark/PySpark, Ray etc.) facilitate it for you out of the box. For example, I stick to Spark/PySpark as a go-to choice to build BigData-scale solutions in the industrial projects. \n\nHowever, launching the infrastructure needed for the platforms like Spark/PySpark requires some resources (RAM, CPU) to be consumed, and it could make it impossible to use on Kaggle kernels with the datasets of a size compared with the data for this contest.",
      "votes": null
    },
    {
      "id": "2643381",
      "postDate": "02/08/2024 19:32:19",
      "content": "<p>thank you for your unique insights!! <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> </p>",
      "rawMarkdown": "thank you for your unique insights!! @gvyshnya",
      "votes": null
    },
    {
      "id": "2643417",
      "postDate": "02/08/2024 20:13:07",
      "content": "<p><a href=\"https://www.kaggle.com/harshalpanchal\" target=\"_blank\">@harshalpanchal</a> : my pleasure, Harshal!</p>",
      "rawMarkdown": "harshalpanchal : my pleasure, Harshal!",
      "votes": null
    },
    {
      "id": "2643489",
      "postDate": "02/08/2024 21:16:05",
      "content": "<p>Great post!<br>\nWhat about the Rapids ecosystem? Speed wise it would be fast as well in loading and preprocessing datasets. One advantage here is that you can make use of a full ecosystem with cudf and cuml. <br>\nWhile polars is an awesome library, tasks like clustering would force a conversion to Pandas first I guess?</p>",
      "rawMarkdown": "Great post!\nWhat about the Rapids ecosystem? Speed wise it would be fast as well in loading and preprocessing datasets. One advantage here is that you can make use of a full ecosystem with cudf and cuml. \nWhile polars is an awesome library, tasks like clustering would force a conversion to Pandas first I guess?",
      "votes": null
    },
    {
      "id": "2644352",
      "postDate": "02/09/2024 12:11:08",
      "content": "<p><a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a>: Hey Thomas, thanks for your kudos!</p>\n<p>Regarding your point on \"tasks like clustering would force a conversion to Pandas first\", it is true. It is even sometimes required for simple visualization needs (since not every graph/chart lib provides polars support out of the box).</p>\n<p>Thank you also for mentioning <em><a href=\"https://rapids.ai/\" target=\"_blank\">rapids</a></em>. I have not had a chance to use it yet. However, brief scanning through its documentation as well as looking through the Kaggle-tailored <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">tutorial</a> displayed a lot of potential in it.</p>",
      "rawMarkdown": "thomasmeiner: Hey Thomas, thanks for your kudos!\n\nRegarding your point on \"tasks like clustering would force a conversion to Pandas first\", it is true. It is even sometimes required for simple visualization needs (since not every graph/chart lib provides polars support out of the box).\n\nThank you also for mentioning *[rapids](https://rapids.ai/)*. I have not had a chance to use it yet. However, brief scanning through its documentation as well as looking through the Kaggle-tailored [tutorial](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets) displayed a lot of potential in it.",
      "votes": null
    },
    {
      "id": "2644834",
      "postDate": "02/09/2024 17:53:49",
      "content": "<p>Thanks for the great tips! Dask can do lazy evaluation so that might help. Also, using functions with limited scope will release memory for local variables once the function is no longer referenced (in theory).</p>",
      "rawMarkdown": "Thanks for the great tips! Dask can do lazy evaluation so that might help. Also, using functions with limited scope will release memory for local variables once the function is no longer referenced (in theory).",
      "votes": null
    },
    {
      "id": "2645007",
      "postDate": "02/09/2024 20:40:45",
      "content": "<p>Thanks for this post. I found Vaex <a href=\"https://pypi.org/project/vaex/\" target=\"_blank\">https://pypi.org/project/vaex/</a> to be a good option for large datasets in the past</p>",
      "rawMarkdown": "Thanks for this post. I found Vaex https://pypi.org/project/vaex/ to be a good option for large datasets in the past",
      "votes": null
    },
    {
      "id": "2645369",
      "postDate": "02/10/2024 07:24:34",
      "content": "<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> : thanks, John! Appreciated your valuable notes on the lazy evaluation in dask as well as on the function-scope-related memory management tricks!</p>",
      "rawMarkdown": "jpmiller : thanks, John! Appreciated your valuable notes on the lazy evaluation in dask as well as on the function-scope-related memory management tricks!",
      "votes": null
    },
    {
      "id": "2645372",
      "postDate": "02/10/2024 07:25:29",
      "content": "<p><a href=\"https://www.kaggle.com/rahi37\" target=\"_blank\">@rahi37</a> : thanks a lot! I have never tried Vaex, and this is a good chance for me to learn something new. Appreciated your knowledge sharing!</p>",
      "rawMarkdown": "rahi37 : thanks a lot! I have never tried Vaex, and this is a good chance for me to learn something new. Appreciated your knowledge sharing!",
      "votes": null
    },
    {
      "id": "2645389",
      "postDate": "02/10/2024 07:41:47",
      "content": "<p>Last time I tested it (a longer time ago) it had hard limitations for group by operations that exceed the usual count, mean etc.</p>",
      "rawMarkdown": "Last time I tested it (a longer time ago) it had hard limitations for group by operations that exceed the usual count, mean etc.",
      "votes": null
    },
    {
      "id": "2645648",
      "postDate": "02/10/2024 11:43:39",
      "content": "<p><a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> : hey Thomas, thank you for sharing your experience with Vaex!</p>",
      "rawMarkdown": "thomasmeiner : hey Thomas, thank you for sharing your experience with Vaex!",
      "votes": null
    },
    {
      "id": "2647042",
      "postDate": "02/11/2024 10:38:19",
      "content": "<p>Thanks a lot dear Georgii <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> ! Too many useful tips! I guess that Your post is most wealth in this competition! Wish You to achieve success in this competition - You deserve it!</p>",
      "rawMarkdown": "Thanks a lot dear Georgii @gvyshnya ! Too many useful tips! I guess that Your post is most wealth in this competition! Wish You to achieve success in this competition - You deserve it!",
      "votes": null
    },
    {
      "id": "2647141",
      "postDate": "02/11/2024 11:31:43",
      "content": "<p><a href=\"https://www.kaggle.com/kapturovalexander\" target=\"_blank\">@kapturovalexander</a> : Hey Alex, thanks a lot!</p>",
      "rawMarkdown": "kapturovalexander : Hey Alex, thanks a lot!",
      "votes": null
    },
    {
      "id": "2670920",
      "postDate": "02/27/2024 07:29:19",
      "content": "<p>Thanks a lot, very useful</p>",
      "rawMarkdown": "Thanks a lot, very useful",
      "votes": null
    },
    {
      "id": "2674473",
      "postDate": "02/29/2024 10:54:52",
      "content": "<p>Hello! Thanks a lot for the great discussion <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> ! It answered some newbie questions I had :-).</p>\n<p>I am trying to understand how to optimize the time and computing restriction of 12 hours in a kaggle notebook:</p>\n<ul>\n<li>A pre-trained model can be used, so we can train our model offline and store it (this is the biggest gain on what can be done, I guess).</li>\n<li>A public dataset can be used: we could think of storing preprocessed data, however, since the test data after submission is not available, we will need to apply the preprocessing to that data online, right?</li>\n</ul>\n<p>So, the time and compute, if completely used, will go into preprocessing test data, and running the inference, is that right? </p>\n<p>EDIT: Another basic question: how about subsampling for EDA? Taking say 10% of case_ids randomly and then only keeping those from </p>",
      "rawMarkdown": "Hello! Thanks a lot for the great discussion @gvyshnya ! It answered some newbie questions I had :-).\n\nI am trying to understand how to optimize the time and computing restriction of 12 hours in a kaggle notebook:\n- A pre-trained model can be used, so we can train our model offline and store it (this is the biggest gain on what can be done, I guess).\n- A public dataset can be used: we could think of storing preprocessed data, however, since the test data after submission is not available, we will need to apply the preprocessing to that data online, right?\n\nSo, the time and compute, if completely used, will go into preprocessing test data, and running the inference, is that right? \n\nEDIT: Another basic question: how about subsampling for EDA? Taking say 10% of case_ids randomly and then only keeping those from",
      "votes": null
    },
    {
      "id": "2674586",
      "postDate": "02/29/2024 12:37:10",
      "content": "<p><a href=\"https://www.kaggle.com/rpicatoste\" target=\"_blank\">@rpicatoste</a> : thanks for your comments.</p>\n<p>Yes, pretrained models  could be one of the performance boosters for the code-based solutions in this competition.</p>\n<p>Regarding the data processing, yes, you will have to apply the  same preprocessing pipelines to both the training data and the testing data. And, yes, it should be the part  of your solution that submits to the competition (in order to handle the unseen part of the testing dataset at the submission time).</p>\n<p>As for subsampling the data in EDA, it could be one of  the strategies to go. However, such an approach (whatever good in theory) has a practical limitation. Your subsample taken for  EDA should have the same distribution of the variables as in the entire training set population in order the EDA to be 100% representative. Otherwise, you will have the drifted insights. So, the bottom line is to be able to extract the fully representative subsample, in order such an EDA strategy to be 100% relevant.</p>\n<p>I hope it is helpful.</p>",
      "rawMarkdown": "rpicatoste : thanks for your comments.\n\nYes, pretrained models  could be one of the performance boosters for the code-based solutions in this competition.\n\nRegarding the data processing, yes, you will have to apply the  same preprocessing pipelines to both the training data and the testing data. And, yes, it should be the part  of your solution that submits to the competition (in order to handle the unseen part of the testing dataset at the submission time).\n\nAs for subsampling the data in EDA, it could be one of  the strategies to go. However, such an approach (whatever good in theory) has a practical limitation. Your subsample taken for  EDA should have the same distribution of the variables as in the entire training set population in order the EDA to be 100% representative. Otherwise, you will have the drifted insights. So, the bottom line is to be able to extract the fully representative subsample, in order such an EDA strategy to be 100% relevant.\n\nI hope it is helpful.",
      "votes": null
    },
    {
      "id": "2740024",
      "postDate": "04/07/2024 13:48:44",
      "content": "<p>Thanks a lot!! This is very helpful</p>",
      "rawMarkdown": "Thanks a lot!! This is very helpful",
      "votes": null
    },
    {
      "id": "2814244",
      "postDate": "05/15/2024 08:23:13",
      "content": "<p>Hello everyone, I am facing \"out of memory\" issues while I submit. </p>\n<p>The training of my model was performed in my pc and I have uploaded it in order to perform only the prediction at Kaggle notebook.</p>\n<p>While the execution is successful the prediction to the hidden test set fails, even though I created a subset of the test set with only 5 rows. I have created lots of features but this doesn't make any sense. </p>\n<p>Has anyone encountered similar problems?</p>",
      "rawMarkdown": "Hello everyone, I am facing \"out of memory\" issues while I submit. \n\nThe training of my model was performed in my pc and I have uploaded it in order to perform only the prediction at Kaggle notebook.\n\nWhile the execution is successful the prediction to the hidden test set fails, even though I created a subset of the test set with only 5 rows. I have created lots of features but this doesn't make any sense. \n\nHas anyone encountered similar problems?",
      "votes": null
    },
    {
      "id": "2817122",
      "postDate": "05/16/2024 18:04:21",
      "content": "<p>try making the notebook on kaggle itself instead of PC.</p>",
      "rawMarkdown": "try making the notebook on kaggle itself instead of PC.",
      "votes": null
    },
    {
      "id": "2817279",
      "postDate": "05/16/2024 19:54:09",
      "content": "<p>Test data files you see are dummy, when you submit your notebook it uses the real (hidden) test set which contains approximately the same number of rows as in the train dataset. So you can test if you can fit the whole train dataset into Kaggle memory, and make predictions on it. Most definitely no, so you should try to reduce memory usage. </p>",
      "rawMarkdown": "Test data files you see are dummy, when you submit your notebook it uses the real (hidden) test set which contains approximately the same number of rows as in the train dataset. So you can test if you can fit the whole train dataset into Kaggle memory, and make predictions on it. Most definitely no, so you should try to reduce memory usage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2643381,
      "author_name": "harshalpanchal",
      "author_url": "",
      "post_date": "02/08/2024 19:32:19",
      "content": "<p>thank you for your unique insights!! <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2643417,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "02/08/2024 20:13:07",
          "content": "<p><a href=\"https://www.kaggle.com/harshalpanchal\" target=\"_blank\">@harshalpanchal</a> : my pleasure, Harshal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2643489,
      "author_name": "thomasmeiner",
      "author_url": "",
      "post_date": "02/08/2024 21:16:05",
      "content": "<p>Great post!<br>\nWhat about the Rapids ecosystem? Speed wise it would be fast as well in loading and preprocessing datasets. One advantage here is that you can make use of a full ecosystem with cudf and cuml. <br>\nWhile polars is an awesome library, tasks like clustering would force a conversion to Pandas first I guess?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2644352,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "02/09/2024 12:11:08",
          "content": "<p><a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a>: Hey Thomas, thanks for your kudos!</p>\n<p>Regarding your point on \"tasks like clustering would force a conversion to Pandas first\", it is true. It is even sometimes required for simple visualization needs (since not every graph/chart lib provides polars support out of the box).</p>\n<p>Thank you also for mentioning <em><a href=\"https://rapids.ai/\" target=\"_blank\">rapids</a></em>. I have not had a chance to use it yet. However, brief scanning through its documentation as well as looking through the Kaggle-tailored <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">tutorial</a> displayed a lot of potential in it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2644834,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "02/09/2024 17:53:49",
      "content": "<p>Thanks for the great tips! Dask can do lazy evaluation so that might help. Also, using functions with limited scope will release memory for local variables once the function is no longer referenced (in theory).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2645369,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "02/10/2024 07:24:34",
          "content": "<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> : thanks, John! Appreciated your valuable notes on the lazy evaluation in dask as well as on the function-scope-related memory management tricks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2645007,
      "author_name": "rahi37",
      "author_url": "",
      "post_date": "02/09/2024 20:40:45",
      "content": "<p>Thanks for this post. I found Vaex <a href=\"https://pypi.org/project/vaex/\" target=\"_blank\">https://pypi.org/project/vaex/</a> to be a good option for large datasets in the past</p>",
      "votes": null,
      "replies": [
        {
          "id": 2645372,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "02/10/2024 07:25:29",
          "content": "<p><a href=\"https://www.kaggle.com/rahi37\" target=\"_blank\">@rahi37</a> : thanks a lot! I have never tried Vaex, and this is a good chance for me to learn something new. Appreciated your knowledge sharing!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2645389,
              "author_name": "thomasmeiner",
              "author_url": "",
              "post_date": "02/10/2024 07:41:47",
              "content": "<p>Last time I tested it (a longer time ago) it had hard limitations for group by operations that exceed the usual count, mean etc.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2645648,
                  "author_name": "gvyshnya",
                  "author_url": "",
                  "post_date": "02/10/2024 11:43:39",
                  "content": "<p><a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> : hey Thomas, thank you for sharing your experience with Vaex!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2647042,
      "author_name": "kapturovalexander",
      "author_url": "",
      "post_date": "02/11/2024 10:38:19",
      "content": "<p>Thanks a lot dear Georgii <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> ! Too many useful tips! I guess that Your post is most wealth in this competition! Wish You to achieve success in this competition - You deserve it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2647141,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "02/11/2024 11:31:43",
          "content": "<p><a href=\"https://www.kaggle.com/kapturovalexander\" target=\"_blank\">@kapturovalexander</a> : Hey Alex, thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2670920,
      "author_name": "ruihaokaggle",
      "author_url": "",
      "post_date": "02/27/2024 07:29:19",
      "content": "<p>Thanks a lot, very useful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2674473,
      "author_name": "rpicatoste",
      "author_url": "",
      "post_date": "02/29/2024 10:54:52",
      "content": "<p>Hello! Thanks a lot for the great discussion <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> ! It answered some newbie questions I had :-).</p>\n<p>I am trying to understand how to optimize the time and computing restriction of 12 hours in a kaggle notebook:</p>\n<ul>\n<li>A pre-trained model can be used, so we can train our model offline and store it (this is the biggest gain on what can be done, I guess).</li>\n<li>A public dataset can be used: we could think of storing preprocessed data, however, since the test data after submission is not available, we will need to apply the preprocessing to that data online, right?</li>\n</ul>\n<p>So, the time and compute, if completely used, will go into preprocessing test data, and running the inference, is that right? </p>\n<p>EDIT: Another basic question: how about subsampling for EDA? Taking say 10% of case_ids randomly and then only keeping those from </p>",
      "votes": null,
      "replies": [
        {
          "id": 2674586,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "02/29/2024 12:37:10",
          "content": "<p><a href=\"https://www.kaggle.com/rpicatoste\" target=\"_blank\">@rpicatoste</a> : thanks for your comments.</p>\n<p>Yes, pretrained models  could be one of the performance boosters for the code-based solutions in this competition.</p>\n<p>Regarding the data processing, yes, you will have to apply the  same preprocessing pipelines to both the training data and the testing data. And, yes, it should be the part  of your solution that submits to the competition (in order to handle the unseen part of the testing dataset at the submission time).</p>\n<p>As for subsampling the data in EDA, it could be one of  the strategies to go. However, such an approach (whatever good in theory) has a practical limitation. Your subsample taken for  EDA should have the same distribution of the variables as in the entire training set population in order the EDA to be 100% representative. Otherwise, you will have the drifted insights. So, the bottom line is to be able to extract the fully representative subsample, in order such an EDA strategy to be 100% relevant.</p>\n<p>I hope it is helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2740024,
      "author_name": "lordpatil",
      "author_url": "",
      "post_date": "04/07/2024 13:48:44",
      "content": "<p>Thanks a lot!! This is very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2814244,
      "author_name": "arispatsopoulos",
      "author_url": "",
      "post_date": "05/15/2024 08:23:13",
      "content": "<p>Hello everyone, I am facing \"out of memory\" issues while I submit. </p>\n<p>The training of my model was performed in my pc and I have uploaded it in order to perform only the prediction at Kaggle notebook.</p>\n<p>While the execution is successful the prediction to the hidden test set fails, even though I created a subset of the test set with only 5 rows. I have created lots of features but this doesn't make any sense. </p>\n<p>Has anyone encountered similar problems?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2817122,
          "author_name": "lordpatil",
          "author_url": "",
          "post_date": "05/16/2024 18:04:21",
          "content": "<p>try making the notebook on kaggle itself instead of PC.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2817279,
          "author_name": "eivolkova",
          "author_url": "",
          "post_date": "05/16/2024 19:54:09",
          "content": "<p>Test data files you see are dummy, when you submit your notebook it uses the real (hidden) test set which contains approximately the same number of rows as in the train dataset. So you can test if you can fit the whole train dataset into Kaggle memory, and make predictions on it. Most definitely no, so you should try to reduce memory usage. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2643121": "# Introduction\n\nThe contest of “Home Credit - Credit Risk Model Stability” imposes us to operate under the significant resource constraints.\n- Handling the large (really BigData-scale) dataset\n- Running it as a code competition on Kaggle notebooks\n\nThe above-mentioned circumstances force us to invent solutions that are\n- Smart in terms of utilizing RAM, CPU/GPU/TPU time, and disk storage. \n- Fast in terms of the execution time (basic Kaggle kernels have the limit of a session to run up to 9 h, with visually reported execution time not exceeding 6 h).\n\nEven if you decide to convert your Kaggle notebook into a Google Collab Pro+ one (it is possible but it comes at a price: https://colab.research.google.com/signup), it would not resolve all of the resource constraint-driven problems for your codebase.\n\nTherefore, the only viable option is to get your head around the best practices of the effective tackling BigData-scale problems in the online resource-limited kernels provided by Kaggle. Some of the best practices are universal, and others are specific to the type of operations you try to accomplish (like EDA-specific or ML Engineering-specific tricks). We are going to review them in the sections below.\n\n# Universal tricks\n\n- Load your data in fast, time-efficient manner (yes, reading BigData with Pandas directly is a pain as well as the drain of the execution time of your session, even if you utilize a batch reading approach); the good example of a fast data loading strategy with *polar* is provided in https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive ; obviously, this is not the only strategy possible to implement – you should be smart to look at the appropriate options (*pyarrrow* to be one of the things to look at, besides *polar*; btw, the nice and fast *pyarrow*-based data loading strategy demonstrated by @tolgadincer in https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet)\n- Optimize the in-memory size of Pandas dataframe (some of the techniques of this sort are shown in https://www.kaggle.com/code/gvyshnya/to-sleep-or-not-to-sleep-deep-eda-dive, for instance)\n- Force regular garbage collection in the  certain points of your notebook, using the capabilities of *gc* package\n- Store the preprocessed training data as a separate Kaggle dataset – it will save the execution of your notebooks doing EDA or/and ML experiments down the road\n\nThe good options for the time-efficient loading could be\n- *polars*\n- *dask*\n- *datatable*\n- novel pandas 2 with pyarrow and schema definition (the good trick of this sort can be reused from the demo in [this notebook](https://www.kaggle.com/code/tolgadincer/sleep-state-fast-data-access-with-parquet) implemented for another contest with BigData-scale data processing requirements)\n- *rapids*\n- etc.\n\n**Note 1:** you can refer to the excellent post [Faster ways to load competition's data](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475389) by @kononenko to see the benchmarking of some of the above-mentioned tools as applied to the data in the dataset for this competition\n\n**Note 2:** Thanks to @thomasmeiner for mentioning *[rapids](https://rapids.ai/)*. Brief scanning through its documentation as well as looking through the Kaggle-tailored [tutorial](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets) displayed a lot of potential in it.\n\n# EDA-specific tricks\n\nPlotting the datasets that are bigger them a RAM memory available to your solution will crash your conventional workflows (like trying to feed a pandas dataframe to a graph lib like plotly or seaborn and then visualize something). \n\nTherefore, more elaborate approaches should be used to achieve your goal. One of the popular patterns of doing it is described below\n- Launch a memory-efficient lightweight database engine (like DuckDB etc.)\n- Mount or import your BigData-scale files into it as tables\n- Pre-calculate the data needed for your charts (histograms, boxplots, bar charts etc.) with SQL queries\n- Visualize your pre-calculated data with a graph library of your choice (with minimal client-side calculations)\n\nThe good case of implementing such a strategy is provided in [Plotting Big Data: Sleeping Series](https://www.kaggle.com/code/gvyshnya/plotting-big-data-sleeping-series) where DuckDB and Plotly were used together to cook the noodles (it is not for this very dataset but the ideas from there could be easily transferred to this data domain).\n\nSQLLite can be also leveraged in such a situation (see [here](https://plotly.com/python/v3/ipython-notebooks/big-data-analytics-with-pandas-and-sqlite/) as an example).\n\n# Specific ML modelling tricks\n\nExercising every general best practice described in the respective section above is essential for the time-efficient ML pipeline built on Kaggle kernels. However, there is one more trick that can help you. \n\nYou can\n- Train your model out of the Kaggle premise\n- Save (serialize) it as a binary file (either pickle or a native TensorFlow model storage format could be used, depending on what a model you train) – in such a way, you get a pre-trained model in your disposal\n- Upload the pre-trained model to a new Kaggle dataset\n- Import such a dataset into your notebook intended to submit predictions to load the pre-trained model from the binary file\n\nOther thing to consider is to unitilize the batch prediction strategy where your model predicts against the subsamples of the test set rather then on the entire test set at once.\n\n# Shall Pandas Be Abandoned?\n\nEven if you become versatile at using BigData-scale platforms like ones mentioned above, you will stilll have to keep *pandas* in your toolbox. Tasks like clustering or visualizing your data would force a conversion of preprocessed data to pandas first. The matter  is, not every special-purpose lib (graphs, clustering, and even ML) provides support for something like polars or dask out of the box.\n\nAt the same time, fast and memory-efficient data preprocessing will make such a niched pandas utilization less painful to your solutions in terms of its memory and CPU footprints.\n\n# Final notes\n\nIf we were not limited to Kaggle kernels in this competition, I would go further to share one more recommendation. It is about taking advantage of **lazy evaluation**. Lazy evaluation refers to the strategy which delays evaluation of an expression until the value is actually needed. Lazy evaluation is an important concept (used especially in functional programming), and if you want to read more about its different usages in Python you can start [here](https://towardsdatascience.com/what-is-lazy-evaluation-in-python-9efb1d3bfed0).\n\nMany industrial platforms (like Spark/PySpark, Ray etc.) facilitate it for you out of the box. For example, I stick to Spark/PySpark as a go-to choice to build BigData-scale solutions in the industrial projects. \n\nHowever, launching the infrastructure needed for the platforms like Spark/PySpark requires some resources (RAM, CPU) to be consumed, and it could make it impossible to use on Kaggle kernels with the datasets of a size compared with the data for this contest.",
    "2643381": "thank you for your unique insights!! @gvyshnya",
    "2643417": "harshalpanchal : my pleasure, Harshal!",
    "2643489": "Great post!\nWhat about the Rapids ecosystem? Speed wise it would be fast as well in loading and preprocessing datasets. One advantage here is that you can make use of a full ecosystem with cudf and cuml. \nWhile polars is an awesome library, tasks like clustering would force a conversion to Pandas first I guess?",
    "2644352": "thomasmeiner: Hey Thomas, thanks for your kudos!\n\nRegarding your point on \"tasks like clustering would force a conversion to Pandas first\", it is true. It is even sometimes required for simple visualization needs (since not every graph/chart lib provides polars support out of the box).\n\nThank you also for mentioning *[rapids](https://rapids.ai/)*. I have not had a chance to use it yet. However, brief scanning through its documentation as well as looking through the Kaggle-tailored [tutorial](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets) displayed a lot of potential in it.",
    "2644834": "Thanks for the great tips! Dask can do lazy evaluation so that might help. Also, using functions with limited scope will release memory for local variables once the function is no longer referenced (in theory).",
    "2645007": "Thanks for this post. I found Vaex https://pypi.org/project/vaex/ to be a good option for large datasets in the past",
    "2645369": "jpmiller : thanks, John! Appreciated your valuable notes on the lazy evaluation in dask as well as on the function-scope-related memory management tricks!",
    "2645372": "rahi37 : thanks a lot! I have never tried Vaex, and this is a good chance for me to learn something new. Appreciated your knowledge sharing!",
    "2645389": "Last time I tested it (a longer time ago) it had hard limitations for group by operations that exceed the usual count, mean etc.",
    "2645648": "thomasmeiner : hey Thomas, thank you for sharing your experience with Vaex!",
    "2647042": "Thanks a lot dear Georgii @gvyshnya ! Too many useful tips! I guess that Your post is most wealth in this competition! Wish You to achieve success in this competition - You deserve it!",
    "2647141": "kapturovalexander : Hey Alex, thanks a lot!",
    "2670920": "Thanks a lot, very useful",
    "2674473": "Hello! Thanks a lot for the great discussion @gvyshnya ! It answered some newbie questions I had :-).\n\nI am trying to understand how to optimize the time and computing restriction of 12 hours in a kaggle notebook:\n- A pre-trained model can be used, so we can train our model offline and store it (this is the biggest gain on what can be done, I guess).\n- A public dataset can be used: we could think of storing preprocessed data, however, since the test data after submission is not available, we will need to apply the preprocessing to that data online, right?\n\nSo, the time and compute, if completely used, will go into preprocessing test data, and running the inference, is that right? \n\nEDIT: Another basic question: how about subsampling for EDA? Taking say 10% of case_ids randomly and then only keeping those from",
    "2674586": "rpicatoste : thanks for your comments.\n\nYes, pretrained models  could be one of the performance boosters for the code-based solutions in this competition.\n\nRegarding the data processing, yes, you will have to apply the  same preprocessing pipelines to both the training data and the testing data. And, yes, it should be the part  of your solution that submits to the competition (in order to handle the unseen part of the testing dataset at the submission time).\n\nAs for subsampling the data in EDA, it could be one of  the strategies to go. However, such an approach (whatever good in theory) has a practical limitation. Your subsample taken for  EDA should have the same distribution of the variables as in the entire training set population in order the EDA to be 100% representative. Otherwise, you will have the drifted insights. So, the bottom line is to be able to extract the fully representative subsample, in order such an EDA strategy to be 100% relevant.\n\nI hope it is helpful.",
    "2740024": "Thanks a lot!! This is very helpful",
    "2814244": "Hello everyone, I am facing \"out of memory\" issues while I submit. \n\nThe training of my model was performed in my pc and I have uploaded it in order to perform only the prediction at Kaggle notebook.\n\nWhile the execution is successful the prediction to the hidden test set fails, even though I created a subset of the test set with only 5 rows. I have created lots of features but this doesn't make any sense. \n\nHas anyone encountered similar problems?",
    "2817122": "try making the notebook on kaggle itself instead of PC.",
    "2817279": "Test data files you see are dummy, when you submit your notebook it uses the real (hidden) test set which contains approximately the same number of rows as in the train dataset. So you can test if you can fit the whole train dataset into Kaggle memory, and make predictions on it. Most definitely no, so you should try to reduce memory usage."
  },
  "source": "meta"
}