{
  "id": 330347,
  "title": "A List of EDA tricks when memory is limited",
  "url": "/competitions/amex-default-prediction/discussion/330347",
  "author_name": "",
  "post_date": "2022-06-11T19:15:03.443948900Z",
  "votes": 56,
  "comment_count": 14,
  "views": 0,
  "content": "<h1>A List of EDA tricks when memory is limited</h1>\n<p>When working with high datasets and common libraries such as Pandas, and trying to plot or visualize parts of the data, we sometimes encounter memory issues simply by trying to plot a simple plot. <br>\nThis is of high importance when doing EDA since fast iteration means everything to data science.</p>\n<p>Here is a short list of tricks and code examples for performing EDA on HUGE datasets. </p>\n<ul>\n<li><strong>Load only a subsample of the data into memory</strong><br>\nWe need to be careful when using subsamples of the data since we are taking the risk of sampling a biased sample. But it's pretty useful early on the EDA round to avoid choking the computer. I have chosen to load the first 10k data points of the dataset.</li>\n</ul>\n<pre><code>df = pd.read_csv(\"data.csv\", header=0, nrows=10000) \n</code></pre>\n<ul>\n<li><strong>Keep your data in compressed format, and decompress on the fly as needed</strong></li>\n</ul>\n<pre><code>import gzip\nf = gzip.open('file.txt.gz', 'r')\nfile_content = f.read()\nf.close()\n</code></pre>\n<ul>\n<li><strong>Store data as text with a compression scheme</strong><br>\nSaving the data as text instead of its original format can be a good choice since we don't work like GP us but with libraries instead. We can opt for saving the data as text files. <br>\nIn the case of large datasets, it's recommended to use a compression algorithm to reduce space.</li>\n</ul>\n<pre><code>import pandas as pd\ndf.to_csv('./file_output.gz', index=False, compression='gzip') \nloaded_data = pd.read_csv('./file_output.gz') \n</code></pre>\n<ul>\n<li><strong>Use numpy's builtin tricks</strong><br>\nUse features supported by numpy to work with dataset that doesn’t fit into memory.<br>\nYou can create an array or matrix that is too big for main memory as an array that is physically stored outside the memory. Being stored outside the main memory, the Numpy array object is an efficient way to limit the memory usage from being filled with the contents of the physical file in its entirety.</li>\n</ul>\n<pre><code>import numpy as np\ndata = np.memmap('memmap', dtype='float32', mode='w+', shape=(10000, 1000000)\n</code></pre>\n<ul>\n<li><strong>Use streaming</strong> methods to do EDA one chank at a time by using <code>chunksize</code> argument of pandas, without loading it all into memory at once</li>\n</ul>\n<pre><code>chunksize = 10**6 # for example\nfor chunk in pd.read_csv('train.csv', chunksize=chunksize):\n    #EDA\n</code></pre>\n<ul>\n<li><p>It’s often a possibility to perform what is needed with much fewer data. For example, a histogram plot can be constructed with a small evenly-spaced sample from the data rather than the entire dataset. Often times it’s unlikely that a plot of every element of the dataset would add much insight.</p></li>\n<li><p><strong>Example:</strong> Aggregating dataset into a smaller number of summary statistics that can be fitted into memory by chunks </p></li>\n</ul>\n<blockquote>\n  <p>Note bessel's correction for the streaming variance </p>\n</blockquote>\n<pre><code>def get_summary_statistics(path):\n    N = 0\n    M2 = 0\n    mean = None\n    variance = None\n    for chunk in pd.read_csv(path):\n        N += len(chunk)\n        for idx, col in enumerate(chunk.columns):\n            data[idx] += chunk[col].values\n            delta = data[idx] - mean\n            mean += (delta/N)\n            M2 += delta*(data[idx] - mean)\n    variance = M2/(N-1)    \n    return mean, variance\n</code></pre>\n<ul>\n<li><strong>Randomly sample</strong> subsets of rows and columns and then run your EDA and plot on this subset to save on memory footprint:</li>\n</ul>\n<pre><code>df_subset = df_subset.sample(frac=0.5, random_state=1)\n</code></pre>\n<ul>\n<li><strong>Load only a subset</strong> of columns each time and run your EDA and plot on this subset to save on memory footprint</li>\n</ul>\n<pre><code>chunk = pd.read_csv(csvfile, usecols=[0,1,2,3,4,5,6,7,8,9])\n</code></pre>\n<ul>\n<li><strong>Load into lighter</strong> numeric formats </li>\n</ul>\n<pre><code>dtype = {'col1': np.int32, 'col2':'float16', 'col3':'float64'}\nchunk = pd.read_csv(csvfile, dtype=dtype)\n</code></pre>\n<ul>\n<li><strong>Don't load duplicates</strong> by first hashing rows chunk by chunk</li>\n</ul>\n<pre><code>for chunk in pd.read_csv(csvfile, chunksize=chunksize):\n    #Calculate the hash of the chunks\n    chunk['hash'] = chunk.apply(lambda x: hash(tuple(x)), axis=1) # any non ram consuming hash function will work here\n    all_chunks_hash.append(chunk)\n    #Keep only unique rows indices\n    unique_rows_idx = all_chunks_hash[all_chunks_hash['hash'].duplicated()==False].index\n    udf = all_chunks_hash.loc[unique_rows_idx]\n</code></pre>\n<blockquote>\n  <p>The largest hurdle to the use of Python for high-performance computing is that it is interpreted, not compiled. This fact is at the heart of every debate about the suitability</p>\n</blockquote>\n<ul>\n<li><strong>For Binary columns</strong> If a dataset has many binary columns, use csr_matrix for a more compact representation</li>\n</ul>\n<pre><code>df_binary_encoded = pd.get_dummies(binary_cols)\ndf_binary_cr_matrix = csr_matrix(df_binary_encoded.values)\ndf_binary_cr_matrix.to_pickle('binary_encoded.pkl')\n</code></pre>\n<ul>\n<li><p><strong>For Category columns</strong> Converting categories to numeric and compressing the numeric representation by converting it to the smallest numeric type possible</p></li>\n<li><p><strong>For Numeric columns</strong> Compressing numeric columns</p></li>\n</ul>\n<blockquote>\n  <p>for example for integers convert to smaller integers</p>\n</blockquote>\n<pre><code>df_numeric['col'] = df_numeric['col'].astype('uint8')\n</code></pre>\n<blockquote>\n  <p>See the famous \"reduce_mem_usage\" for more on downsampled data types</p>\n</blockquote>\n<ul>\n<li><p><strong>Use memory-efficient built-in data structures</strong></p></li>\n<li><p>Discard prior history as soon as we can and always use <code>del</code> to delete old objects and then collect memory with <code>gc.collect()</code></p></li>\n</ul>\n<pre><code>del old_object\ngc.collect()\n</code></pre>\n<ul>\n<li>In General, try not to use visualization methods that require the sequence length to be known beforehand or visualization methods that are \"fancy\" in such cases that they are not \"a must\". </li>\n</ul>\n<p><strong>Debugging Memory Issues</strong></p>\n<ul>\n<li>Use the Python <code>memory_profiler</code> library to see where memory is getting eaten up</li>\n</ul>\n<pre><code>from memory_profiler import memory_usage\nres = memory_usage(proc())\nindex = res.index(max(res))\nwith pd.open_csv('myfile.csv') as df:\n    for dataset in df:\n        proc(dataset)\n</code></pre>\n<ul>\n<li>For example, is it eating memory up with your EDA and visualization code, or is it being eaten up in the loading process? Or is the file being stored on disk in an inefficient format (sparsed areas in the data take up just as much space as sparse areas in the data)</li>\n</ul>\n<p><strong>Bonus Section: Bypassing Memory Limits with swapfiles</strong></p>\n<p>Common limits that we encounter:</p>\n<ul>\n<li><p>memory limits set by the OS (such as those used with virtual machines).</p></li>\n<li><p>Some applications limit how much memory can be used by a given process.</p></li>\n<li><p>Python’s multiprocessing module will fail to fail with an OSError error if one attempts to launch a process that exceeds the amount of memory available.</p></li>\n<li><p>Memory profile typically shows that most of the memory is used by object allocations and only a fraction is used for allocated data. Traditionally languages such as C have placed memory usage limits on the size of data that can be allocated, but this is receding as allocations may be performed out of the physical address space</p></li>\n<li><p>Swapfiles on Linux-based systems (swapfiles are files that can act as a backup memory for our RAM memory)</p></li>\n</ul>\n<p>To set up swapfile and extend our system's ram we simply can use the following commands (Example is 4GB swapfile on path: \"/swapfile\":</p>\n<pre><code>sudo fallocate -l 4G /swapfile \nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n</code></pre>\n<p><strong>Bonus: Multiprocessing with limited memory</strong></p>\n<ul>\n<li>Use joblib to use multiple cores while the dataset is still loaded on disk.</li>\n<li>Use <strong>multiprocessing.Pool</strong> to load the data and distribute it between cores. </li>\n</ul>\n<p><strong>Note:</strong> Here is a comparison between the two approaches:</p>\n<ul>\n<li><strong><code>import joblib</code></strong> vs <strong><code>import multiprocessing</code></strong>: The multiprocessing module relies on spawning processes and passing messages between them, while the joblib module uses functions provided by the operating system. Both enable the parallel execution of tasks within a single computer and on a single shared-memory system, but they support different patterns of use. The choice of which to use often breaks down to the pattern of use (batch or online) and whether the data is shared between tasks.</li>\n</ul>\n<p><strong>Bonus: GPU memory tricks</strong><br>\nGPUs are not created equal and so have different capabilities in terms of the amount of memory they can access. <br>\nTo save up on GPU ram footprint we can use one of the following approaches:</p>\n<ul>\n<li><strong>Batching</strong> loading data in smaller (cpu) batches to the gpu and in parallel loading next batch from disk to avoid waiting on the gpu with no data to process.</li>\n<li><strong>Downsample</strong> the data to a smaller size</li>\n<li><strong>Compress</strong> the data to a size that will fit on the gpu.</li>\n</ul>\n<p>To improve GPU ram usage, we use wrappers that replace CPU-based training routines with GPU-based equivalents, for example: <strong>cupy vs numpy</strong>:</p>\n<pre><code>import cupy as cp\n\nx = cp.asarray(x)\ny = cp.asarray(y)\n</code></pre>\n<p><strong>Other Tricks</strong></p>\n<ul>\n<li><p><strong>Fancy Plots:</strong> For other kinds of plots that require loading all the data, downsampling can be useful to both allow plots to be generated faster than if the entire dataset were used and can also provide a means to view all of the datasets at once. Generating the best-looking plots will require some experimentation but some general tips include:</p></li>\n<li><p><strong>Resampling</strong> to a lower resolution on images and time series can reduce the dataset size by orders of magnitude.</p></li>\n<li><p><strong>Visualizing/Loading Top/Bottom Results By a Model</strong> If loading the entire dataset at once won’t fit in memory, we can draw plots of just the top results and then color/shade the other data points in a way that indicates that the rest of the data are less important.</p></li>\n<li><p><strong>Plotting Library:</strong> Changing to a different plotting library more suitable for large datasets some libraries provide plotting methods that are more memory efficient than others. If generating the plots takes too long and/or won’t fit in memory, a different plotting library might simply \"just work\" making the same type of plot.</p></li>\n<li><p><strong>Streaming numeric</strong> data from files does have some disadvantages in the event that we need a sorted copy of the data, for example. To address this need, it is possible to store the data in memory but use a different strategy for storing the data in order to reduce the amount of memory required. Python’s built-in sorted and <code>heapq</code> modules can be used to sort data in a streaming fashion that requires less memory overhead.</p></li>\n<li><p><strong>Distribution</strong> of the dataset can be summarized by histograms without first loading the entire dataset.<br>\ngenerate_histograms_from_data_stream() provides an example of the code, which generates the histograms one chunk at a time without requiring any data to be loaded into memory at once.</p></li>\n<li><p><strong>Just reduce_mem_usage it:</strong> Losing precision by storing numbers as bytes or even bits can dramatically reduce memory requirements at the expense of precision, in terms of EDA: this \"might\" be OK. still: be careful. </p></li>\n</ul>",
  "messages": [
    {
      "id": "1817773",
      "postDate": "06/11/2022 19:15:03",
      "content": "<h1>A List of EDA tricks when memory is limited</h1>\n<p>When working with high datasets and common libraries such as Pandas, and trying to plot or visualize parts of the data, we sometimes encounter memory issues simply by trying to plot a simple plot. <br>\nThis is of high importance when doing EDA since fast iteration means everything to data science.</p>\n<p>Here is a short list of tricks and code examples for performing EDA on HUGE datasets. </p>\n<ul>\n<li><strong>Load only a subsample of the data into memory</strong><br>\nWe need to be careful when using subsamples of the data since we are taking the risk of sampling a biased sample. But it's pretty useful early on the EDA round to avoid choking the computer. I have chosen to load the first 10k data points of the dataset.</li>\n</ul>\n<pre><code>df = pd.read_csv(\"data.csv\", header=0, nrows=10000) \n</code></pre>\n<ul>\n<li><strong>Keep your data in compressed format, and decompress on the fly as needed</strong></li>\n</ul>\n<pre><code>import gzip\nf = gzip.open('file.txt.gz', 'r')\nfile_content = f.read()\nf.close()\n</code></pre>\n<ul>\n<li><strong>Store data as text with a compression scheme</strong><br>\nSaving the data as text instead of its original format can be a good choice since we don't work like GP us but with libraries instead. We can opt for saving the data as text files. <br>\nIn the case of large datasets, it's recommended to use a compression algorithm to reduce space.</li>\n</ul>\n<pre><code>import pandas as pd\ndf.to_csv('./file_output.gz', index=False, compression='gzip') \nloaded_data = pd.read_csv('./file_output.gz') \n</code></pre>\n<ul>\n<li><strong>Use numpy's builtin tricks</strong><br>\nUse features supported by numpy to work with dataset that doesn’t fit into memory.<br>\nYou can create an array or matrix that is too big for main memory as an array that is physically stored outside the memory. Being stored outside the main memory, the Numpy array object is an efficient way to limit the memory usage from being filled with the contents of the physical file in its entirety.</li>\n</ul>\n<pre><code>import numpy as np\ndata = np.memmap('memmap', dtype='float32', mode='w+', shape=(10000, 1000000)\n</code></pre>\n<ul>\n<li><strong>Use streaming</strong> methods to do EDA one chank at a time by using <code>chunksize</code> argument of pandas, without loading it all into memory at once</li>\n</ul>\n<pre><code>chunksize = 10**6 # for example\nfor chunk in pd.read_csv('train.csv', chunksize=chunksize):\n    #EDA\n</code></pre>\n<ul>\n<li><p>It’s often a possibility to perform what is needed with much fewer data. For example, a histogram plot can be constructed with a small evenly-spaced sample from the data rather than the entire dataset. Often times it’s unlikely that a plot of every element of the dataset would add much insight.</p></li>\n<li><p><strong>Example:</strong> Aggregating dataset into a smaller number of summary statistics that can be fitted into memory by chunks </p></li>\n</ul>\n<blockquote>\n  <p>Note bessel's correction for the streaming variance </p>\n</blockquote>\n<pre><code>def get_summary_statistics(path):\n    N = 0\n    M2 = 0\n    mean = None\n    variance = None\n    for chunk in pd.read_csv(path):\n        N += len(chunk)\n        for idx, col in enumerate(chunk.columns):\n            data[idx] += chunk[col].values\n            delta = data[idx] - mean\n            mean += (delta/N)\n            M2 += delta*(data[idx] - mean)\n    variance = M2/(N-1)    \n    return mean, variance\n</code></pre>\n<ul>\n<li><strong>Randomly sample</strong> subsets of rows and columns and then run your EDA and plot on this subset to save on memory footprint:</li>\n</ul>\n<pre><code>df_subset = df_subset.sample(frac=0.5, random_state=1)\n</code></pre>\n<ul>\n<li><strong>Load only a subset</strong> of columns each time and run your EDA and plot on this subset to save on memory footprint</li>\n</ul>\n<pre><code>chunk = pd.read_csv(csvfile, usecols=[0,1,2,3,4,5,6,7,8,9])\n</code></pre>\n<ul>\n<li><strong>Load into lighter</strong> numeric formats </li>\n</ul>\n<pre><code>dtype = {'col1': np.int32, 'col2':'float16', 'col3':'float64'}\nchunk = pd.read_csv(csvfile, dtype=dtype)\n</code></pre>\n<ul>\n<li><strong>Don't load duplicates</strong> by first hashing rows chunk by chunk</li>\n</ul>\n<pre><code>for chunk in pd.read_csv(csvfile, chunksize=chunksize):\n    #Calculate the hash of the chunks\n    chunk['hash'] = chunk.apply(lambda x: hash(tuple(x)), axis=1) # any non ram consuming hash function will work here\n    all_chunks_hash.append(chunk)\n    #Keep only unique rows indices\n    unique_rows_idx = all_chunks_hash[all_chunks_hash['hash'].duplicated()==False].index\n    udf = all_chunks_hash.loc[unique_rows_idx]\n</code></pre>\n<blockquote>\n  <p>The largest hurdle to the use of Python for high-performance computing is that it is interpreted, not compiled. This fact is at the heart of every debate about the suitability</p>\n</blockquote>\n<ul>\n<li><strong>For Binary columns</strong> If a dataset has many binary columns, use csr_matrix for a more compact representation</li>\n</ul>\n<pre><code>df_binary_encoded = pd.get_dummies(binary_cols)\ndf_binary_cr_matrix = csr_matrix(df_binary_encoded.values)\ndf_binary_cr_matrix.to_pickle('binary_encoded.pkl')\n</code></pre>\n<ul>\n<li><p><strong>For Category columns</strong> Converting categories to numeric and compressing the numeric representation by converting it to the smallest numeric type possible</p></li>\n<li><p><strong>For Numeric columns</strong> Compressing numeric columns</p></li>\n</ul>\n<blockquote>\n  <p>for example for integers convert to smaller integers</p>\n</blockquote>\n<pre><code>df_numeric['col'] = df_numeric['col'].astype('uint8')\n</code></pre>\n<blockquote>\n  <p>See the famous \"reduce_mem_usage\" for more on downsampled data types</p>\n</blockquote>\n<ul>\n<li><p><strong>Use memory-efficient built-in data structures</strong></p></li>\n<li><p>Discard prior history as soon as we can and always use <code>del</code> to delete old objects and then collect memory with <code>gc.collect()</code></p></li>\n</ul>\n<pre><code>del old_object\ngc.collect()\n</code></pre>\n<ul>\n<li>In General, try not to use visualization methods that require the sequence length to be known beforehand or visualization methods that are \"fancy\" in such cases that they are not \"a must\". </li>\n</ul>\n<p><strong>Debugging Memory Issues</strong></p>\n<ul>\n<li>Use the Python <code>memory_profiler</code> library to see where memory is getting eaten up</li>\n</ul>\n<pre><code>from memory_profiler import memory_usage\nres = memory_usage(proc())\nindex = res.index(max(res))\nwith pd.open_csv('myfile.csv') as df:\n    for dataset in df:\n        proc(dataset)\n</code></pre>\n<ul>\n<li>For example, is it eating memory up with your EDA and visualization code, or is it being eaten up in the loading process? Or is the file being stored on disk in an inefficient format (sparsed areas in the data take up just as much space as sparse areas in the data)</li>\n</ul>\n<p><strong>Bonus Section: Bypassing Memory Limits with swapfiles</strong></p>\n<p>Common limits that we encounter:</p>\n<ul>\n<li><p>memory limits set by the OS (such as those used with virtual machines).</p></li>\n<li><p>Some applications limit how much memory can be used by a given process.</p></li>\n<li><p>Python’s multiprocessing module will fail to fail with an OSError error if one attempts to launch a process that exceeds the amount of memory available.</p></li>\n<li><p>Memory profile typically shows that most of the memory is used by object allocations and only a fraction is used for allocated data. Traditionally languages such as C have placed memory usage limits on the size of data that can be allocated, but this is receding as allocations may be performed out of the physical address space</p></li>\n<li><p>Swapfiles on Linux-based systems (swapfiles are files that can act as a backup memory for our RAM memory)</p></li>\n</ul>\n<p>To set up swapfile and extend our system's ram we simply can use the following commands (Example is 4GB swapfile on path: \"/swapfile\":</p>\n<pre><code>sudo fallocate -l 4G /swapfile \nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n</code></pre>\n<p><strong>Bonus: Multiprocessing with limited memory</strong></p>\n<ul>\n<li>Use joblib to use multiple cores while the dataset is still loaded on disk.</li>\n<li>Use <strong>multiprocessing.Pool</strong> to load the data and distribute it between cores. </li>\n</ul>\n<p><strong>Note:</strong> Here is a comparison between the two approaches:</p>\n<ul>\n<li><strong><code>import joblib</code></strong> vs <strong><code>import multiprocessing</code></strong>: The multiprocessing module relies on spawning processes and passing messages between them, while the joblib module uses functions provided by the operating system. Both enable the parallel execution of tasks within a single computer and on a single shared-memory system, but they support different patterns of use. The choice of which to use often breaks down to the pattern of use (batch or online) and whether the data is shared between tasks.</li>\n</ul>\n<p><strong>Bonus: GPU memory tricks</strong><br>\nGPUs are not created equal and so have different capabilities in terms of the amount of memory they can access. <br>\nTo save up on GPU ram footprint we can use one of the following approaches:</p>\n<ul>\n<li><strong>Batching</strong> loading data in smaller (cpu) batches to the gpu and in parallel loading next batch from disk to avoid waiting on the gpu with no data to process.</li>\n<li><strong>Downsample</strong> the data to a smaller size</li>\n<li><strong>Compress</strong> the data to a size that will fit on the gpu.</li>\n</ul>\n<p>To improve GPU ram usage, we use wrappers that replace CPU-based training routines with GPU-based equivalents, for example: <strong>cupy vs numpy</strong>:</p>\n<pre><code>import cupy as cp\n\nx = cp.asarray(x)\ny = cp.asarray(y)\n</code></pre>\n<p><strong>Other Tricks</strong></p>\n<ul>\n<li><p><strong>Fancy Plots:</strong> For other kinds of plots that require loading all the data, downsampling can be useful to both allow plots to be generated faster than if the entire dataset were used and can also provide a means to view all of the datasets at once. Generating the best-looking plots will require some experimentation but some general tips include:</p></li>\n<li><p><strong>Resampling</strong> to a lower resolution on images and time series can reduce the dataset size by orders of magnitude.</p></li>\n<li><p><strong>Visualizing/Loading Top/Bottom Results By a Model</strong> If loading the entire dataset at once won’t fit in memory, we can draw plots of just the top results and then color/shade the other data points in a way that indicates that the rest of the data are less important.</p></li>\n<li><p><strong>Plotting Library:</strong> Changing to a different plotting library more suitable for large datasets some libraries provide plotting methods that are more memory efficient than others. If generating the plots takes too long and/or won’t fit in memory, a different plotting library might simply \"just work\" making the same type of plot.</p></li>\n<li><p><strong>Streaming numeric</strong> data from files does have some disadvantages in the event that we need a sorted copy of the data, for example. To address this need, it is possible to store the data in memory but use a different strategy for storing the data in order to reduce the amount of memory required. Python’s built-in sorted and <code>heapq</code> modules can be used to sort data in a streaming fashion that requires less memory overhead.</p></li>\n<li><p><strong>Distribution</strong> of the dataset can be summarized by histograms without first loading the entire dataset.<br>\ngenerate_histograms_from_data_stream() provides an example of the code, which generates the histograms one chunk at a time without requiring any data to be loaded into memory at once.</p></li>\n<li><p><strong>Just reduce_mem_usage it:</strong> Losing precision by storing numbers as bytes or even bits can dramatically reduce memory requirements at the expense of precision, in terms of EDA: this \"might\" be OK. still: be careful. </p></li>\n</ul>",
      "rawMarkdown": "# A List of EDA tricks when memory is limited\n\nWhen working with high datasets and common libraries such as Pandas, and trying to plot or visualize parts of the data, we sometimes encounter memory issues simply by trying to plot a simple plot. \nThis is of high importance when doing EDA since fast iteration means everything to data science.\n\nHere is a short list of tricks and code examples for performing EDA on HUGE datasets. \n\n- **Load only a subsample of the data into memory**\nWe need to be careful when using subsamples of the data since we are taking the risk of sampling a biased sample. But it's pretty useful early on the EDA round to avoid choking the computer. I have chosen to load the first 10k data points of the dataset.\n\n```python\ndf = pd.read_csv(\"data.csv\", header=0, nrows=10000) \n```\n- **Keep your data in compressed format, and decompress on the fly as needed**\n\n```python\nimport gzip\nf = gzip.open('file.txt.gz', 'r')\nfile_content = f.read()\nf.close()\n```\n- **Store data as text with a compression scheme**\nSaving the data as text instead of its original format can be a good choice since we don't work like GP us but with libraries instead. We can opt for saving the data as text files. \nIn the case of large datasets, it's recommended to use a compression algorithm to reduce space.\n\n```python\nimport pandas as pd\ndf.to_csv('./file_output.gz', index=False, compression='gzip') \nloaded_data = pd.read_csv('./file_output.gz') \n```\n\n- **Use numpy's builtin tricks**\nUse features supported by numpy to work with dataset that doesn’t fit into memory.\nYou can create an array or matrix that is too big for main memory as an array that is physically stored outside the memory. Being stored outside the main memory, the Numpy array object is an efficient way to limit the memory usage from being filled with the contents of the physical file in its entirety.\n\n```python\nimport numpy as np\ndata = np.memmap('memmap', dtype='float32', mode='w+', shape=(10000, 1000000)\n```\n\n- **Use streaming** methods to do EDA one chank at a time by using `chunksize` argument of pandas, without loading it all into memory at once\n\n```python\nchunksize = 10**6 # for example\nfor chunk in pd.read_csv('train.csv', chunksize=chunksize):\n    #EDA\n```\n\n- It’s often a possibility to perform what is needed with much fewer data. For example, a histogram plot can be constructed with a small evenly-spaced sample from the data rather than the entire dataset. Often times it’s unlikely that a plot of every element of the dataset would add much insight.\n\n\n- **Example:** Aggregating dataset into a smaller number of summary statistics that can be fitted into memory by chunks \n\n> Note bessel's correction for the streaming variance \n\n```python\ndef get_summary_statistics(path):\n    N = 0\n    M2 = 0\n    mean = None\n    variance = None\n    for chunk in pd.read_csv(path):\n        N += len(chunk)\n        for idx, col in enumerate(chunk.columns):\n            data[idx] += chunk[col].values\n            delta = data[idx] - mean\n            mean += (delta/N)\n            M2 += delta*(data[idx] - mean)\n    variance = M2/(N-1)    \n    return mean, variance\n```\n\n- **Randomly sample** subsets of rows and columns and then run your EDA and plot on this subset to save on memory footprint:\n\n```python\ndf_subset = df_subset.sample(frac=0.5, random_state=1)\n```\n\n- **Load only a subset** of columns each time and run your EDA and plot on this subset to save on memory footprint\n\n```python\nchunk = pd.read_csv(csvfile, usecols=[0,1,2,3,4,5,6,7,8,9])\n```\n\n- **Load into lighter** numeric formats \n```python\ndtype = {'col1': np.int32, 'col2':'float16', 'col3':'float64'}\nchunk = pd.read_csv(csvfile, dtype=dtype)\n```\n\n- **Don't load duplicates** by first hashing rows chunk by chunk\n\n```python\nfor chunk in pd.read_csv(csvfile, chunksize=chunksize):\n    #Calculate the hash of the chunks\n    chunk['hash'] = chunk.apply(lambda x: hash(tuple(x)), axis=1) # any non ram consuming hash function will work here\n    all_chunks_hash.append(chunk)\n    #Keep only unique rows indices\n    unique_rows_idx = all_chunks_hash[all_chunks_hash['hash'].duplicated()==False].index\n    udf = all_chunks_hash.loc[unique_rows_idx]\n```\n\n> The largest hurdle to the use of Python for high-performance computing is that it is interpreted, not compiled. This fact is at the heart of every debate about the suitability\n\n- **For Binary columns** If a dataset has many binary columns, use csr_matrix for a more compact representation\n\n```python\ndf_binary_encoded = pd.get_dummies(binary_cols)\ndf_binary_cr_matrix = csr_matrix(df_binary_encoded.values)\ndf_binary_cr_matrix.to_pickle('binary_encoded.pkl')\n```\n\n- **For Category columns** Converting categories to numeric and compressing the numeric representation by converting it to the smallest numeric type possible\n\n- **For Numeric columns** Compressing numeric columns\n\n> for example for integers convert to smaller integers\n\n```python\ndf_numeric['col'] = df_numeric['col'].astype('uint8')\n```\n> See the famous \"reduce_mem_usage\" for more on downsampled data types\n\n- **Use memory-efficient built-in data structures**\n\n- Discard prior history as soon as we can and always use `del` to delete old objects and then collect memory with `gc.collect()`\n\n```python\ndel old_object\ngc.collect()\n```\n\n- In General, try not to use visualization methods that require the sequence length to be known beforehand or visualization methods that are \"fancy\" in such cases that they are not \"a must\". \n\n\n**Debugging Memory Issues**\n\n- Use the Python `memory_profiler` library to see where memory is getting eaten up\n```python\nfrom memory_profiler import memory_usage\nres = memory_usage(proc())\nindex = res.index(max(res))\nwith pd.open_csv('myfile.csv') as df:\n    for dataset in df:\n        proc(dataset)\n```\n\n- For example, is it eating memory up with your EDA and visualization code, or is it being eaten up in the loading process? Or is the file being stored on disk in an inefficient format (sparsed areas in the data take up just as much space as sparse areas in the data)\n\n\n\n**Bonus Section: Bypassing Memory Limits with swapfiles**\n\nCommon limits that we encounter:\n- memory limits set by the OS (such as those used with virtual machines).\n- Some applications limit how much memory can be used by a given process.\n- Python’s multiprocessing module will fail to fail with an OSError error if one attempts to launch a process that exceeds the amount of memory available.\n\n- Memory profile typically shows that most of the memory is used by object allocations and only a fraction is used for allocated data. Traditionally languages such as C have placed memory usage limits on the size of data that can be allocated, but this is receding as allocations may be performed out of the physical address space\n\n- Swapfiles on Linux-based systems (swapfiles are files that can act as a backup memory for our RAM memory)\n\nTo set up swapfile and extend our system's ram we simply can use the following commands (Example is 4GB swapfile on path: \"/swapfile\":\n```\nsudo fallocate -l 4G /swapfile \nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n```\n\n**Bonus: Multiprocessing with limited memory**\n\n- Use joblib to use multiple cores while the dataset is still loaded on disk.\n- Use **multiprocessing.Pool** to load the data and distribute it between cores. \n\n**Note:** Here is a comparison between the two approaches:\n\n- **`import joblib`** vs **`import multiprocessing`**: The multiprocessing module relies on spawning processes and passing messages between them, while the joblib module uses functions provided by the operating system. Both enable the parallel execution of tasks within a single computer and on a single shared-memory system, but they support different patterns of use. The choice of which to use often breaks down to the pattern of use (batch or online) and whether the data is shared between tasks.\n\n**Bonus: GPU memory tricks**\nGPUs are not created equal and so have different capabilities in terms of the amount of memory they can access. \nTo save up on GPU ram footprint we can use one of the following approaches:\n- **Batching** loading data in smaller (cpu) batches to the gpu and in parallel loading next batch from disk to avoid waiting on the gpu with no data to process.\n- **Downsample** the data to a smaller size\n- **Compress** the data to a size that will fit on the gpu.\n\nTo improve GPU ram usage, we use wrappers that replace CPU-based training routines with GPU-based equivalents, for example: **cupy vs numpy**:\n```python\nimport cupy as cp\n\nx = cp.asarray(x)\ny = cp.asarray(y)\n\n```\n\n**Other Tricks**\n\n- **Fancy Plots:** For other kinds of plots that require loading all the data, downsampling can be useful to both allow plots to be generated faster than if the entire dataset were used and can also provide a means to view all of the datasets at once. Generating the best-looking plots will require some experimentation but some general tips include:\n\n- **Resampling** to a lower resolution on images and time series can reduce the dataset size by orders of magnitude.\n\n- **Visualizing/Loading Top/Bottom Results By a Model** If loading the entire dataset at once won’t fit in memory, we can draw plots of just the top results and then color/shade the other data points in a way that indicates that the rest of the data are less important.\n\n- **Plotting Library:** Changing to a different plotting library more suitable for large datasets some libraries provide plotting methods that are more memory efficient than others. If generating the plots takes too long and/or won’t fit in memory, a different plotting library might simply \"just work\" making the same type of plot.\n\n- **Streaming numeric** data from files does have some disadvantages in the event that we need a sorted copy of the data, for example. To address this need, it is possible to store the data in memory but use a different strategy for storing the data in order to reduce the amount of memory required. Python’s built-in sorted and `heapq` modules can be used to sort data in a streaming fashion that requires less memory overhead.\n\n- **Distribution** of the dataset can be summarized by histograms without first loading the entire dataset.\ngenerate_histograms_from_data_stream() provides an example of the code, which generates the histograms one chunk at a time without requiring any data to be loaded into memory at once.\n\n- **Just reduce_mem_usage it:** Losing precision by storing numbers as bytes or even bits can dramatically reduce memory requirements at the expense of precision, in terms of EDA: this \"might\" be OK. still: be careful.",
      "votes": null
    },
    {
      "id": "1817776",
      "postDate": "06/11/2022 19:22:39",
      "content": "<p>Very useful tricks. Thanks for sharing</p>",
      "rawMarkdown": "Very useful tricks. Thanks for sharing",
      "votes": null
    },
    {
      "id": "1817783",
      "postDate": "06/11/2022 19:37:28",
      "content": "<p>This is one of the best and most informative posts I have seen on this platform lately.</p>",
      "rawMarkdown": "This is one of the best and most informative posts I have seen on this platform lately.",
      "votes": null
    },
    {
      "id": "1817788",
      "postDate": "06/11/2022 19:51:32",
      "content": "<p>Thank you! I'm glad I was able to help :)</p>",
      "rawMarkdown": "Thank you! I'm glad I was able to help :)",
      "votes": null
    },
    {
      "id": "1817790",
      "postDate": "06/11/2022 19:52:10",
      "content": "<p>Thanks! If you got any more you know, I will be more then happy to know about them! </p>",
      "rawMarkdown": "Thanks! If you got any more you know, I will be more then happy to know about them!",
      "votes": null
    },
    {
      "id": "1819329",
      "postDate": "06/13/2022 16:43:06",
      "content": "<p>Nice! Didn't understand half of them, but now I have something to work on ;)</p>",
      "rawMarkdown": "Nice! Didn't understand half of them, but now I have something to work on ;)",
      "votes": null
    },
    {
      "id": "1819860",
      "postDate": "06/14/2022 07:31:36",
      "content": "<p>This should be a blog post <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> ✨<br>\nSuper useful for competitions that have large datasets!</p>",
      "rawMarkdown": "This should be a blog post @thedevastator ✨\nSuper useful for competitions that have large datasets!",
      "votes": null
    },
    {
      "id": "1819876",
      "postDate": "06/14/2022 07:55:46",
      "content": "<p>Great insights on the challenges that working with large datasets can present. Really liked the ideas around optimal memory usage techniques. A grasp of the message provided by a sample of the data can help with the overall story. </p>",
      "rawMarkdown": "Great insights on the challenges that working with large datasets can present. Really liked the ideas around optimal memory usage techniques. A grasp of the message provided by a sample of the data can help with the overall story.",
      "votes": null
    },
    {
      "id": "1820288",
      "postDate": "06/14/2022 13:57:12",
      "content": "<p>Amazing, thanks for sharing !</p>",
      "rawMarkdown": "Amazing, thanks for sharing !",
      "votes": null
    },
    {
      "id": "1820364",
      "postDate": "06/14/2022 15:05:35",
      "content": "<p>Thanks for sharing. It is so useful! Have a nice Data Trip!</p>",
      "rawMarkdown": "Thanks for sharing. It is so useful! Have a nice Data Trip!",
      "votes": null
    },
    {
      "id": "1820556",
      "postDate": "06/14/2022 18:31:40",
      "content": "<p>Amazing explanations to handle big datasets. Love it.<br>\nMaybe you can also use datatable for loading the data faster. I made a notebook for this one. <br>\n<a href=\"https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster\" target=\"_blank\">Loading your data Faster</a></p>",
      "rawMarkdown": "Amazing explanations to handle big datasets. Love it.\nMaybe you can also use datatable for loading the data faster. I made a notebook for this one. \n[Loading your data Faster](https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster)",
      "votes": null
    },
    {
      "id": "1821425",
      "postDate": "06/15/2022 14:30:57",
      "content": "<p>Very Helpful Tips !!  Thanks for sharing :)</p>",
      "rawMarkdown": "Very Helpful Tips !!  Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1821443",
      "postDate": "06/15/2022 14:38:19",
      "content": "<p>Nice tricks <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> and congratulations on becoming <strong>kaggle expert</strong>. Very impressive the way your going forward </p>",
      "rawMarkdown": "Nice tricks @thedevastator and congratulations on becoming **kaggle expert**. Very impressive the way your going forward",
      "votes": null
    },
    {
      "id": "1821481",
      "postDate": "06/15/2022 15:10:37",
      "content": "<p>Amazing tip, thanks!</p>",
      "rawMarkdown": "Amazing tip, thanks!",
      "votes": null
    },
    {
      "id": "1906573",
      "postDate": "08/20/2022 02:16:35",
      "content": "<p>Most helpful tips!! I'm faced Memory issue, but it might solve my error, I think.😃👍</p>",
      "rawMarkdown": "Most helpful tips!! I'm faced Memory issue, but it might solve my error, I think.😃👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1817776,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "06/11/2022 19:22:39",
      "content": "<p>Very useful tricks. Thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1817790,
          "author_name": "thedevastator",
          "author_url": "",
          "post_date": "06/11/2022 19:52:10",
          "content": "<p>Thanks! If you got any more you know, I will be more then happy to know about them! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1817783,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/11/2022 19:37:28",
      "content": "<p>This is one of the best and most informative posts I have seen on this platform lately.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1817788,
          "author_name": "thedevastator",
          "author_url": "",
          "post_date": "06/11/2022 19:51:32",
          "content": "<p>Thank you! I'm glad I was able to help :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1819329,
      "author_name": "alonsourbano",
      "author_url": "",
      "post_date": "06/13/2022 16:43:06",
      "content": "<p>Nice! Didn't understand half of them, but now I have something to work on ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1819860,
      "author_name": "ruchi798",
      "author_url": "",
      "post_date": "06/14/2022 07:31:36",
      "content": "<p>This should be a blog post <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> ✨<br>\nSuper useful for competitions that have large datasets!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1819876,
      "author_name": "datajmcn",
      "author_url": "",
      "post_date": "06/14/2022 07:55:46",
      "content": "<p>Great insights on the challenges that working with large datasets can present. Really liked the ideas around optimal memory usage techniques. A grasp of the message provided by a sample of the data can help with the overall story. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1820288,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "06/14/2022 13:57:12",
      "content": "<p>Amazing, thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1820364,
      "author_name": "sikymddeberias",
      "author_url": "",
      "post_date": "06/14/2022 15:05:35",
      "content": "<p>Thanks for sharing. It is so useful! Have a nice Data Trip!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1820556,
      "author_name": "naiborhujosua",
      "author_url": "",
      "post_date": "06/14/2022 18:31:40",
      "content": "<p>Amazing explanations to handle big datasets. Love it.<br>\nMaybe you can also use datatable for loading the data faster. I made a notebook for this one. <br>\n<a href=\"https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster\" target=\"_blank\">Loading your data Faster</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821425,
      "author_name": "gaurav126",
      "author_url": "",
      "post_date": "06/15/2022 14:30:57",
      "content": "<p>Very Helpful Tips !!  Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821443,
      "author_name": "thamotharan",
      "author_url": "",
      "post_date": "06/15/2022 14:38:19",
      "content": "<p>Nice tricks <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> and congratulations on becoming <strong>kaggle expert</strong>. Very impressive the way your going forward </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821481,
      "author_name": "knedliky",
      "author_url": "",
      "post_date": "06/15/2022 15:10:37",
      "content": "<p>Amazing tip, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1906573,
      "author_name": "jackwilliams3rd",
      "author_url": "",
      "post_date": "08/20/2022 02:16:35",
      "content": "<p>Most helpful tips!! I'm faced Memory issue, but it might solve my error, I think.😃👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1817773": "# A List of EDA tricks when memory is limited\n\nWhen working with high datasets and common libraries such as Pandas, and trying to plot or visualize parts of the data, we sometimes encounter memory issues simply by trying to plot a simple plot. \nThis is of high importance when doing EDA since fast iteration means everything to data science.\n\nHere is a short list of tricks and code examples for performing EDA on HUGE datasets. \n\n- **Load only a subsample of the data into memory**\nWe need to be careful when using subsamples of the data since we are taking the risk of sampling a biased sample. But it's pretty useful early on the EDA round to avoid choking the computer. I have chosen to load the first 10k data points of the dataset.\n\n```python\ndf = pd.read_csv(\"data.csv\", header=0, nrows=10000) \n```\n- **Keep your data in compressed format, and decompress on the fly as needed**\n\n```python\nimport gzip\nf = gzip.open('file.txt.gz', 'r')\nfile_content = f.read()\nf.close()\n```\n- **Store data as text with a compression scheme**\nSaving the data as text instead of its original format can be a good choice since we don't work like GP us but with libraries instead. We can opt for saving the data as text files. \nIn the case of large datasets, it's recommended to use a compression algorithm to reduce space.\n\n```python\nimport pandas as pd\ndf.to_csv('./file_output.gz', index=False, compression='gzip') \nloaded_data = pd.read_csv('./file_output.gz') \n```\n\n- **Use numpy's builtin tricks**\nUse features supported by numpy to work with dataset that doesn’t fit into memory.\nYou can create an array or matrix that is too big for main memory as an array that is physically stored outside the memory. Being stored outside the main memory, the Numpy array object is an efficient way to limit the memory usage from being filled with the contents of the physical file in its entirety.\n\n```python\nimport numpy as np\ndata = np.memmap('memmap', dtype='float32', mode='w+', shape=(10000, 1000000)\n```\n\n- **Use streaming** methods to do EDA one chank at a time by using `chunksize` argument of pandas, without loading it all into memory at once\n\n```python\nchunksize = 10**6 # for example\nfor chunk in pd.read_csv('train.csv', chunksize=chunksize):\n    #EDA\n```\n\n- It’s often a possibility to perform what is needed with much fewer data. For example, a histogram plot can be constructed with a small evenly-spaced sample from the data rather than the entire dataset. Often times it’s unlikely that a plot of every element of the dataset would add much insight.\n\n\n- **Example:** Aggregating dataset into a smaller number of summary statistics that can be fitted into memory by chunks \n\n> Note bessel's correction for the streaming variance \n\n```python\ndef get_summary_statistics(path):\n    N = 0\n    M2 = 0\n    mean = None\n    variance = None\n    for chunk in pd.read_csv(path):\n        N += len(chunk)\n        for idx, col in enumerate(chunk.columns):\n            data[idx] += chunk[col].values\n            delta = data[idx] - mean\n            mean += (delta/N)\n            M2 += delta*(data[idx] - mean)\n    variance = M2/(N-1)    \n    return mean, variance\n```\n\n- **Randomly sample** subsets of rows and columns and then run your EDA and plot on this subset to save on memory footprint:\n\n```python\ndf_subset = df_subset.sample(frac=0.5, random_state=1)\n```\n\n- **Load only a subset** of columns each time and run your EDA and plot on this subset to save on memory footprint\n\n```python\nchunk = pd.read_csv(csvfile, usecols=[0,1,2,3,4,5,6,7,8,9])\n```\n\n- **Load into lighter** numeric formats \n```python\ndtype = {'col1': np.int32, 'col2':'float16', 'col3':'float64'}\nchunk = pd.read_csv(csvfile, dtype=dtype)\n```\n\n- **Don't load duplicates** by first hashing rows chunk by chunk\n\n```python\nfor chunk in pd.read_csv(csvfile, chunksize=chunksize):\n    #Calculate the hash of the chunks\n    chunk['hash'] = chunk.apply(lambda x: hash(tuple(x)), axis=1) # any non ram consuming hash function will work here\n    all_chunks_hash.append(chunk)\n    #Keep only unique rows indices\n    unique_rows_idx = all_chunks_hash[all_chunks_hash['hash'].duplicated()==False].index\n    udf = all_chunks_hash.loc[unique_rows_idx]\n```\n\n> The largest hurdle to the use of Python for high-performance computing is that it is interpreted, not compiled. This fact is at the heart of every debate about the suitability\n\n- **For Binary columns** If a dataset has many binary columns, use csr_matrix for a more compact representation\n\n```python\ndf_binary_encoded = pd.get_dummies(binary_cols)\ndf_binary_cr_matrix = csr_matrix(df_binary_encoded.values)\ndf_binary_cr_matrix.to_pickle('binary_encoded.pkl')\n```\n\n- **For Category columns** Converting categories to numeric and compressing the numeric representation by converting it to the smallest numeric type possible\n\n- **For Numeric columns** Compressing numeric columns\n\n> for example for integers convert to smaller integers\n\n```python\ndf_numeric['col'] = df_numeric['col'].astype('uint8')\n```\n> See the famous \"reduce_mem_usage\" for more on downsampled data types\n\n- **Use memory-efficient built-in data structures**\n\n- Discard prior history as soon as we can and always use `del` to delete old objects and then collect memory with `gc.collect()`\n\n```python\ndel old_object\ngc.collect()\n```\n\n- In General, try not to use visualization methods that require the sequence length to be known beforehand or visualization methods that are \"fancy\" in such cases that they are not \"a must\". \n\n\n**Debugging Memory Issues**\n\n- Use the Python `memory_profiler` library to see where memory is getting eaten up\n```python\nfrom memory_profiler import memory_usage\nres = memory_usage(proc())\nindex = res.index(max(res))\nwith pd.open_csv('myfile.csv') as df:\n    for dataset in df:\n        proc(dataset)\n```\n\n- For example, is it eating memory up with your EDA and visualization code, or is it being eaten up in the loading process? Or is the file being stored on disk in an inefficient format (sparsed areas in the data take up just as much space as sparse areas in the data)\n\n\n\n**Bonus Section: Bypassing Memory Limits with swapfiles**\n\nCommon limits that we encounter:\n- memory limits set by the OS (such as those used with virtual machines).\n- Some applications limit how much memory can be used by a given process.\n- Python’s multiprocessing module will fail to fail with an OSError error if one attempts to launch a process that exceeds the amount of memory available.\n\n- Memory profile typically shows that most of the memory is used by object allocations and only a fraction is used for allocated data. Traditionally languages such as C have placed memory usage limits on the size of data that can be allocated, but this is receding as allocations may be performed out of the physical address space\n\n- Swapfiles on Linux-based systems (swapfiles are files that can act as a backup memory for our RAM memory)\n\nTo set up swapfile and extend our system's ram we simply can use the following commands (Example is 4GB swapfile on path: \"/swapfile\":\n```\nsudo fallocate -l 4G /swapfile \nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n```\n\n**Bonus: Multiprocessing with limited memory**\n\n- Use joblib to use multiple cores while the dataset is still loaded on disk.\n- Use **multiprocessing.Pool** to load the data and distribute it between cores. \n\n**Note:** Here is a comparison between the two approaches:\n\n- **`import joblib`** vs **`import multiprocessing`**: The multiprocessing module relies on spawning processes and passing messages between them, while the joblib module uses functions provided by the operating system. Both enable the parallel execution of tasks within a single computer and on a single shared-memory system, but they support different patterns of use. The choice of which to use often breaks down to the pattern of use (batch or online) and whether the data is shared between tasks.\n\n**Bonus: GPU memory tricks**\nGPUs are not created equal and so have different capabilities in terms of the amount of memory they can access. \nTo save up on GPU ram footprint we can use one of the following approaches:\n- **Batching** loading data in smaller (cpu) batches to the gpu and in parallel loading next batch from disk to avoid waiting on the gpu with no data to process.\n- **Downsample** the data to a smaller size\n- **Compress** the data to a size that will fit on the gpu.\n\nTo improve GPU ram usage, we use wrappers that replace CPU-based training routines with GPU-based equivalents, for example: **cupy vs numpy**:\n```python\nimport cupy as cp\n\nx = cp.asarray(x)\ny = cp.asarray(y)\n\n```\n\n**Other Tricks**\n\n- **Fancy Plots:** For other kinds of plots that require loading all the data, downsampling can be useful to both allow plots to be generated faster than if the entire dataset were used and can also provide a means to view all of the datasets at once. Generating the best-looking plots will require some experimentation but some general tips include:\n\n- **Resampling** to a lower resolution on images and time series can reduce the dataset size by orders of magnitude.\n\n- **Visualizing/Loading Top/Bottom Results By a Model** If loading the entire dataset at once won’t fit in memory, we can draw plots of just the top results and then color/shade the other data points in a way that indicates that the rest of the data are less important.\n\n- **Plotting Library:** Changing to a different plotting library more suitable for large datasets some libraries provide plotting methods that are more memory efficient than others. If generating the plots takes too long and/or won’t fit in memory, a different plotting library might simply \"just work\" making the same type of plot.\n\n- **Streaming numeric** data from files does have some disadvantages in the event that we need a sorted copy of the data, for example. To address this need, it is possible to store the data in memory but use a different strategy for storing the data in order to reduce the amount of memory required. Python’s built-in sorted and `heapq` modules can be used to sort data in a streaming fashion that requires less memory overhead.\n\n- **Distribution** of the dataset can be summarized by histograms without first loading the entire dataset.\ngenerate_histograms_from_data_stream() provides an example of the code, which generates the histograms one chunk at a time without requiring any data to be loaded into memory at once.\n\n- **Just reduce_mem_usage it:** Losing precision by storing numbers as bytes or even bits can dramatically reduce memory requirements at the expense of precision, in terms of EDA: this \"might\" be OK. still: be careful.",
    "1817776": "Very useful tricks. Thanks for sharing",
    "1817783": "This is one of the best and most informative posts I have seen on this platform lately.",
    "1817788": "Thank you! I'm glad I was able to help :)",
    "1817790": "Thanks! If you got any more you know, I will be more then happy to know about them!",
    "1819329": "Nice! Didn't understand half of them, but now I have something to work on ;)",
    "1819860": "This should be a blog post @thedevastator ✨\nSuper useful for competitions that have large datasets!",
    "1819876": "Great insights on the challenges that working with large datasets can present. Really liked the ideas around optimal memory usage techniques. A grasp of the message provided by a sample of the data can help with the overall story.",
    "1820288": "Amazing, thanks for sharing !",
    "1820364": "Thanks for sharing. It is so useful! Have a nice Data Trip!",
    "1820556": "Amazing explanations to handle big datasets. Love it.\nMaybe you can also use datatable for loading the data faster. I made a notebook for this one. \n[Loading your data Faster](https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster)",
    "1821425": "Very Helpful Tips !!  Thanks for sharing :)",
    "1821443": "Nice tricks @thedevastator and congratulations on becoming **kaggle expert**. Very impressive the way your going forward",
    "1821481": "Amazing tip, thanks!",
    "1906573": "Most helpful tips!! I'm faced Memory issue, but it might solve my error, I think.😃👍"
  },
  "source": "meta"
}