{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Reading Parquet Files - RAM/CPU Optimization \n\nThe Bengali AI dataset is used to explore the different methods available for reading Parquet files (pandas + pyarrow).\n\nA common source of trouble for Kernel Only Compeitions, is Out-Of-Memory errors, and the 120 minute submission time limit.\n\nThis notebook contains:\n- Syntax and performance for reading Parquet via both Pandas and Pyarrow\n- Kaggle Kernel RAM/CPU allocation\n  - 18G RAM\n  - 2x Intel(R) Xeon(R) CPU @ 2.00GHz CPU\n- RAM optimized generator function around `pandas.read_parquet()`\n  - trade 50% RAM (1700MB -> 780MB) for 2x disk IO time (5.8min -> 10.2min runtime) \n- RAM/CPU profiling of implict dataframe dtype casting \n  - beware of implicit cast between `unit8` -> `float64` = 8x memory usage\n  - `skimage.measure.block_reduce(train, (1,2,2,1), func=np.mean, cval=0)` can downsample images\n","metadata":{}},{"cell_type":"markdown","source":"# RAM/CPU Available In Kaggle Kernel\n\nIn theory there is 18GB of Kaggle RAM, but loading the entire dataset at once often causes out of memory errors, and doesn't leave anything for the tensorflow model. In practice, datasets need to be loaded one file at a time (or even 75% of a file) to permit a successful compile and submission run.","metadata":{}},{"cell_type":"code","source":"!free -h","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:40.527548Z","iopub.execute_input":"2022-06-05T06:25:40.529652Z","iopub.status.idle":"2022-06-05T06:25:41.349002Z","shell.execute_reply.started":"2022-06-05T06:25:40.529382Z","shell.execute_reply":"2022-06-05T06:25:41.347744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2x Intel(R) Xeon(R) CPU @ 2.00GHz CPU\n\nIn theory this might allow for optimizations using `pathos.multiprocessing`","metadata":{}},{"cell_type":"code","source":"!cat /proc/cpuinfo","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:41.351093Z","iopub.execute_input":"2022-06-05T06:25:41.352075Z","iopub.status.idle":"2022-06-05T06:25:42.120988Z","shell.execute_reply.started":"2022-06-05T06:25:41.352033Z","shell.execute_reply":"2022-06-05T06:25:42.119415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Available Libaries","metadata":{}},{"cell_type":"markdown","source":"Both `pandas` and `pyarrow` are the two possible libaries to use\n\nNOTE: `parquet` and `fastparquet` are not in the Kaggle pip repo, even with the latest available docket images. Whilst these can be obtained via `!pip install parquet fastparquet`, this requires an internet connection which is not allowed for Kernel only competitions.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport pyarrow\nimport pyarrow.parquet as pq\nfrom pyarrow.parquet import ParquetFile\n\ntry:   import parquet\nexcept Exception as exception: print(exception)\n    \ntry:   import fastparquet\nexcept Exception as exception: print(exception)    ","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:42.123371Z","iopub.execute_input":"2022-06-05T06:25:42.123913Z","iopub.status.idle":"2022-06-05T06:25:42.132855Z","shell.execute_reply.started":"2022-06-05T06:25:42.123856Z","shell.execute_reply":"2022-06-05T06:25:42.131853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Other imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport pyarrow\nimport gc\nimport time\nimport glob\nimport sys\nimport humanize\nimport math\nimport psutil\nimport gc\nimport simplejson\nimport skimage\nimport skimage.measure\nfrom timeit import timeit\nfrom time import sleep\nfrom pyarrow.parquet import ParquetFile\nimport pyarrow\nimport pyarrow.parquet as pq\nimport signal\nfrom contextlib import contextmanager\n\npd.set_option('display.max_columns',   500)\npd.set_option('display.max_colwidth',  None)","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:42.135661Z","iopub.execute_input":"2022-06-05T06:25:42.136348Z","iopub.status.idle":"2022-06-05T06:25:43.252536Z","shell.execute_reply.started":"2022-06-05T06:25:42.136299Z","shell.execute_reply":"2022-06-05T06:25:43.251787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"@contextmanager\ndef timeout(time):\n    # Register a function to raise a TimeoutError on the signal.\n    signal.signal(signal.SIGALRM, raise_timeout)\n    # Schedule the signal to be sent after ``time``.\n    signal.alarm(time)\n\n    try:\n        yield\n    except TimeoutError:\n        pass\n    finally:\n        # Unregister the signal so it won't be triggered\n        # if the timeout is not reached.\n        signal.signal(signal.SIGALRM, signal.SIG_IGN)\n\ndef raise_timeout(signum, frame):\n    raise TimeoutError","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:43.253584Z","iopub.execute_input":"2022-06-05T06:25:43.254571Z","iopub.status.idle":"2022-06-05T06:25:43.26093Z","shell.execute_reply.started":"2022-06-05T06:25:43.254531Z","shell.execute_reply":"2022-06-05T06:25:43.260052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Memory Profiler Decorator\nIts also worth mentioning the memory_profiler `@profile` decorator for interactive debugging. \n- NOTE: @profile / %mprun can only be used on functions defined in physical files, and not in the IPython environment.\n- https://pypi.org/project/memory-profiler/","metadata":{}},{"cell_type":"code","source":"from memory_profiler import profile","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:43.262202Z","iopub.execute_input":"2022-06-05T06:25:43.263285Z","iopub.status.idle":"2022-06-05T06:25:43.284059Z","shell.execute_reply.started":"2022-06-05T06:25:43.263238Z","shell.execute_reply":"2022-06-05T06:25:43.282788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read Parquet via Pandas\n\nPandas is the simplest and recommended option\n- it takes 40s seconds to physically read all the data\n- pandas dataset is 6.5GB in RAM. \n","metadata":{"trusted":true}},{"cell_type":"code","source":"!python --version  # Python 3.6.6 :: Anaconda, Inc == original (2020-03-14)\n                   # Python 3.6.6 :: Anaconda, Inc == latest   (2022-06-05)","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:43.285419Z","iopub.execute_input":"2022-06-05T06:25:43.2859Z","iopub.status.idle":"2022-06-05T06:25:44.077054Z","shell.execute_reply.started":"2022-06-05T06:25:43.285866Z","shell.execute_reply":"2022-06-05T06:25:44.075583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.__version__  # 0.25.3 == original (2020-03-14)\n                # 1.3.5  == latest   (2022-06-05)","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:44.079339Z","iopub.execute_input":"2022-06-05T06:25:44.07993Z","iopub.status.idle":"2022-06-05T06:25:44.092975Z","shell.execute_reply.started":"2022-06-05T06:25:44.079875Z","shell.execute_reply":"2022-06-05T06:25:44.091847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filenames = sorted(glob.glob('../input/bengaliai-cv19/train_image_data_*.parquet')); filenames","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:44.095121Z","iopub.execute_input":"2022-06-05T06:25:44.095667Z","iopub.status.idle":"2022-06-05T06:25:44.110051Z","shell.execute_reply.started":"2022-06-05T06:25:44.095619Z","shell.execute_reply":"2022-06-05T06:25:44.108708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_parquet_via_pandas(files=4, cast='uint8', resize=1):\n    gc.collect(); sleep(5);  # wait for gc to complete\n    memory_before = psutil.virtual_memory()[3]\n    # NOTE: loading all the files into a list variable, then applying pd.concat() into a second variable, uses double the memory\n    df = pd.concat([ \n        pd.read_parquet(filename).set_index('image_id', drop=True).astype('uint8')\n        for filename in filenames[:files] \n    ])\n    memory_end= psutil.virtual_memory()[3]        \n\n    print( \"  sys.getsizeof():\", humanize.naturalsize(sys.getsizeof(df)) )\n    print( \"  memory total:   \", humanize.naturalsize(memory_end - memory_before), '+system', humanize.naturalsize(memory_before) )        \n    return df\n\n\ngc.collect(); sleep(2);  # wait for gc to complete\nprint('single file:')\ntime_start = time.time()\nread_parquet_via_pandas(files=1); gc.collect()\nprint(f\"  time:            {time.time() - time_start:.1f}s\" )\nprint('------------------------------')\nprint('pd.concat() all files:')\ntime_start = time.time()\nread_parquet_via_pandas(files=4); gc.collect()\nprint(f\"  time:            {time.time() - time_start:.1f}s\" )\npass","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:25:44.113148Z","iopub.execute_input":"2022-06-05T06:25:44.113531Z","iopub.status.idle":"2022-06-05T06:26:32.876961Z","shell.execute_reply.started":"2022-06-05T06:25:44.113497Z","shell.execute_reply":"2022-06-05T06:26:32.875655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read Parquet via PyArrow","metadata":{"trusted":true}},{"cell_type":"markdown","source":"Creating a `ParquetFile` is very quick, and memory efficent. It only creates a pointer to the file, but allows us to read the metadata.\n\nHowever there the dataset only contains a single `row_group`, meaning the file can only be read out as a single chunk (no easy row-by-row streaming)","metadata":{}},{"cell_type":"code","source":"import pyarrow\nimport pyarrow.parquet as pq\nfrom pyarrow.parquet import ParquetFile\n\npyarrow.__version__  # original 0.16.0 == (2020-03-14)\n                     # latest.  8.0.0  == (2022-06-05)","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:26:32.877892Z","iopub.status.idle":"2022-06-05T06:26:32.878598Z","shell.execute_reply.started":"2022-06-05T06:26:32.878397Z","shell.execute_reply":"2022-06-05T06:26:32.878418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# DOCS: https://arrow.apache.org/docs/python/generated/pyarrow.parquet.ParquetFile.html\ndef read_parquet_via_pyarrow_file():\n    pqfiles = [ ParquetFile(filename) for filename in filenames ]\n    print( \"sys.getsizeof\", humanize.naturalsize(sys.getsizeof(pqfiles)) )\n    for pqfile in pqfiles[0:1]: print(pqfile.metadata)\n    return pqfiles\n\ngc.collect(); sleep(2);  # wait for gc to complete\ntime_start = time.time()\nread_parquet_via_pyarrow_file(); gc.collect()\nprint( f\"time: {time.time() - time_start:.1f}s\" )\npass","metadata":{"_cell_guid":"","_uuid":"","execution":{"iopub.status.busy":"2022-06-05T06:26:32.879529Z","iopub.status.idle":"2022-06-05T06:26:32.879891Z","shell.execute_reply.started":"2022-06-05T06:26:32.879718Z","shell.execute_reply":"2022-06-05T06:26:32.879735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using a pyarrow.Table is faster than pandas (`28s` vs `45s`), but uses more memory (`7.6GB` vs `6.5GB`) and causes an Out-Of-Memory exception if everything is read at once","metadata":{}},{"cell_type":"code","source":"# DOCS: https://arrow.apache.org/docs/python/parquet.html\n# DOCS: https://arrow.apache.org/docs/python/generated/pyarrow.Table.html\n# NOTE: Attempting to read all tables into memory, causes an out of memory exception\ndef read_parquet_via_pyarrow_table():\n    shapes  = []\n    classes = []\n    sizes   = 0\n    for filename in filenames:\n        table = pq.read_table(filename) \n        shapes.append( table.shape )\n        classes.append( table.__class__ )\n        size = sys.getsizeof(table); sizes += size\n        print(\"sys.getsizeof(): \",   humanize.naturalsize(sys.getsizeof(table))  )        \n    print(\"sys.getsizeof() total:\", humanize.naturalsize(sizes) )\n    print(\"classes:\", classes)\n    print(\"shapes: \",  shapes)    \n\n\ngc.collect(); sleep(2);  # wait for gc to complete\ntime_start = time.time()\nread_parquet_via_pyarrow_table(); gc.collect()\nprint( f\"time:   {time.time() - time_start:.1f}s\" )\npass","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:26:32.881094Z","iopub.status.idle":"2022-06-05T06:26:32.881523Z","shell.execute_reply.started":"2022-06-05T06:26:32.881316Z","shell.execute_reply":"2022-06-05T06:26:32.881334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A generator can be written around pyarrow, but this still reads the contents of an entire file into memory and this function is really slow","metadata":{}},{"cell_type":"code","source":"import time, psutil, gc\n\ngc.collect(); sleep(2)  # wait for gc to complete\nmem_before   = psutil.virtual_memory()[3]\nmemory_usage = []\n\ndef read_parquet_via_pyarrow_table_generator(batch_size=128):\n    for filename in filenames[0:1]:  # only loop over one file for demonstration purposes\n        gc.collect(); sleep(1)\n        for batch in pq.read_table(filename).to_batches(batch_size):\n            mem_current = psutil.virtual_memory()[3]\n            memory_usage.append( mem_current - mem_before )\n            yield batch.to_pandas()\n\n\ntime_start = time.time()\ncount = 0\nfor batch in read_parquet_via_pyarrow_table_generator():\n    count += 1\n\nprint( \"time:             \", time.time() - time_start )\nprint( \"count:            \", count )\nprint( \"min memory_usage: \", humanize.naturalsize(min(memory_usage))  )\nprint( \"max memory_usage: \", humanize.naturalsize(max(memory_usage))  )\nprint( \"avg memory_usage: \", humanize.naturalsize(np.mean(memory_usage)) )\npass    ","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:26:32.882532Z","iopub.status.idle":"2022-06-05T06:26:32.88289Z","shell.execute_reply.started":"2022-06-05T06:26:32.882718Z","shell.execute_reply":"2022-06-05T06:26:32.882734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Pandas Batch Generator Function","metadata":{}},{"cell_type":"markdown","source":"It is possible to write a batch generator using pandas. In theory this should save memory, at the expense of disk IO. \n\n- Timer show that disk IO increase linarly with the number of filesystem reads. \n- Memory measurements require `gc.collect(); sleep(1)`, but show that average/min memory reduces linearly with filesystem reads\n\nThere are 8 files to read (including test files in the submission), so the tradeoffs are as follows:\n- reads_per_file 1 |  44s * 8 =  5.8min + 1700MB RAM (potentually crashing the kernel)\n- reads_per_file 2 |  77s * 8 = 10.2min +  781MB RAM (minimum required to solve the memory bottleneck)\n- reads_per_file 3 | 112s * 8 = 14.9min +  508MB RAM (1/8th of total 120min runtime)\n- reads_per_file 5 | 183s * 8 = 24.4min +  314MB RAM (1/5th of total 120min runtime)\n\nThis is a memory/time tradeoff, but demonstrates a practical solution to out-of-memory errors","metadata":{}},{"cell_type":"code","source":"memory_before = psutil.virtual_memory()[3]\nmemory_usage  = []\n\ndef read_parquet_via_pandas_generator(batch_size=128, reads_per_file=5):\n    for filename in filenames:\n        num_rows    = ParquetFile(filename).metadata.num_rows\n        cache_size  = math.ceil( num_rows / batch_size / reads_per_file ) * batch_size\n        batch_count = math.ceil( cache_size / batch_size )\n        for n_read in range(reads_per_file):\n            cache = pd.read_parquet(filename).iloc[ cache_size * n_read : cache_size * (n_read+1) ].copy()\n            gc.collect(); sleep(1);  # sleep(1) is required to allow measurement of the garbage collector\n            for n_batch in range(batch_count):            \n                memory_current = psutil.virtual_memory()[3]\n                memory_usage.append( memory_current - memory_before )                \n                yield cache[ batch_size * n_batch : batch_size * (n_batch+1) ].copy()\n\n                \nfor reads_per_file in [1,2,3,5]: \n    gc.collect(); sleep(5);  # wait for gc to complete\n    memory_before = psutil.virtual_memory()[3]\n    memory_usage  = []\n    \n    time_start = time.time()\n    count = 0\n    for batch in read_parquet_via_pandas_generator(batch_size=128, reads_per_file=reads_per_file):\n        count += 1\n        \n    print( \"reads_per_file\", reads_per_file, '|', \n           'time', int(time.time() - time_start),'s', '|', \n           'count', count,  '|',\n           'memory', {\n                \"min\": humanize.naturalsize(min(memory_usage)),\n                \"max\": humanize.naturalsize(max(memory_usage)),\n                \"avg\": humanize.naturalsize(np.mean(memory_usage)),\n                \"+system\": humanize.naturalsize(memory_before),               \n            }\n    )\npass    ","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:26:32.883828Z","iopub.status.idle":"2022-06-05T06:26:32.884174Z","shell.execute_reply.started":"2022-06-05T06:26:32.884012Z","shell.execute_reply":"2022-06-05T06:26:32.884029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dtypes and Memory Usage","metadata":{}},{"cell_type":"markdown","source":"Memory useage can vary by an order of magnitude based on the implcit cast dtype. \n\n- Raw pixel values are read from the parquet file as `uint8`\n- `/ 255.0` or `skimage.measure.block_reduce()` will do an implict cast of `int` -> `float64`\n- `float64` results in a datastructure 8x as large as `uint8` (`13.0 GB` vs `1.8 GB`)\n  - This can be avoided by doing an explict cast to `float16` (`3.3 GB`)\n- `skimage.measure.block_reduce(df, (1,n,n,1), func=np.mean, cval=0)` == `AveragePooling2D(n)` \n  - reduces data structure memory by `n^2` \n\nCPU time: \n- `float32` (+0.5s) is the fastest cast; `float16` (+8s) is 2x slower than cast `float64` (+4s).\n- `skimage.measure.block_reduce()` is an expensive operation (3-5x IO read time)","metadata":{}},{"cell_type":"code","source":"def read_single_parquet_via_pandas_with_cast(dtype='uint8', normalize=False, denoise=False, invert=True, resize=1, resize_fn=None):\n    gc.collect(); sleep(2);\n    \n    memory_before = psutil.virtual_memory()[3]\n    time_start = time.time()        \n    \n    train = (pd.read_parquet(filenames[0])\n               .set_index('image_id', drop=True)\n               .values.astype(dtype)\n               .reshape(-1, 137, 236, 1))\n    \n    if invert:                                         # Colors | 0 = black      | 255 = white\n        train = (255-train)                            # invert | 0 = background | 255 = line\n   \n    if denoise:                                        # Set small pixel values to background 0\n        if invert: train *= (train >= 25)              #   0 = background | 255 = line  | np.mean() == 12\n        else:      train += (255-train)*(train >= 230) # 255 = background |   0 = line  | np.mean() == 244     \n        \n    if isinstance(resize, bool) and resize == True:\n        resize = 2    # Reduce image size by 2x\n    if resize and resize != 1:                  \n        # NOTEBOOK: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n        # Out of the different resize functions:\n        # - np.mean(dtype=uint8) produces produces fragmented images (needs float16 to work properly - but RAM intensive)\n        # - np.median() produces the most accurate downsampling\n        # - np.max() produces an enhanced image with thicker lines (maybe slightly easier to read)\n        # - np.min() produces a  dehanced image with thiner lines (harder to read)\n        resize_fn = resize_fn or (np.max if invert else np.min)\n        cval      = 0 if invert else 255\n        train = skimage.measure.block_reduce(train, (1, resize,resize, 1), cval=cval, func=resize_fn)  # train.shape = (50210, 137, 236, 1)\n        \n    if normalize:\n        train = train / 255.0          # division casts: int -> float64 \n\n\n    time_end     = time.time()\n    memory_after = psutil.virtual_memory()[3] \n    return ( \n        str(round(time_end - time_start,2)).rjust(5),\n        # str(sys.getsizeof(train)),\n        str(memory_after - memory_before).rjust(5), \n        str(train.shape).ljust(20),\n        str(train.dtype).ljust(7),\n    )\n\n\n# for dtype in ['uint8', 'uint16', 'uint32', 'float16', 'float32']:  # float64                    cause OOM error (2020-03-14)\nfor dtype in ['uint8', 'uint16', 'float16', ]:                       # float64 + uint32 + float32 cause OOM error (2022-06-05)\n    seconds, memory, shape, dtype = read_single_parquet_via_pandas_with_cast(dtype=dtype)\n    print(f'dtype {dtype}'.ljust(18) + f'| {dtype} | {shape} | {seconds}s | {humanize.naturalsize(memory).rjust(8)}')\n\nfor denoise in [False, True]:\n    seconds, memory, shape, dtype = read_single_parquet_via_pandas_with_cast(denoise=denoise)\n    print(f'denoise {denoise}'.ljust(18) + f'| {dtype} | {shape} | {seconds}s | {humanize.naturalsize(memory).rjust(8)}')\n\n# for normalize in [False, True]:  # True + False work    (2020-03-14)\nfor normalize in [False]:          # True cause OOM error (2022-06-05)\n    seconds, memory, shape, dtype = read_single_parquet_via_pandas_with_cast(normalize=normalize)\n    print(f'normalize {normalize}'.ljust(18) + f'| {dtype} | {shape} | {seconds}s | {humanize.naturalsize(memory).rjust(8)}')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:26:46.258237Z","iopub.execute_input":"2022-06-05T06:26:46.259001Z","iopub.status.idle":"2022-06-05T06:28:37.081858Z","shell.execute_reply.started":"2022-06-05T06:26:46.258944Z","shell.execute_reply":"2022-06-05T06:28:37.080657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# division casts: int -> float64 \n# for dtype in ['float16', 'float32', 'float64']:  # float64 causes OOM error (2020-03-14)\n# for dtype in ['float16', 'float32']:             # float32 causes OOM error (2022-06-05)\nfor dtype in ['float16']:             \n    seconds, memory, shape, dtype = read_single_parquet_via_pandas_with_cast(dtype=dtype, normalize=True)\n    print(f'normalize {dtype}'.ljust(18) + f'| {dtype} | {shape} | {seconds}s | {humanize.naturalsize(memory).rjust(8)}')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:28:37.083664Z","iopub.execute_input":"2022-06-05T06:28:37.084043Z","iopub.status.idle":"2022-06-05T06:30:01.863183Z","shell.execute_reply.started":"2022-06-05T06:28:37.08401Z","shell.execute_reply":"2022-06-05T06:30:01.862069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# skimage.measure.block_reduce() casts: unit8 -> float64    \nfor resize in [2, 3, 4]:\n    # for dtype in ['uint8', 'float16', 'float32']:  # 'float32' almosts causes OOM error (2020-03-14)\n    for dtype in ['uint8', 'float16']:               # 'float32' now     causes OOM error (2022-06-05)      \n        gc.collect()\n        with timeout(10*60):\n            seconds, memory, shape, dtype = read_single_parquet_via_pandas_with_cast(dtype=dtype, resize=resize)\n            print(f'resize {resize} {dtype}'.ljust(18) + f'| {dtype} | {shape} | {seconds}s | {humanize.naturalsize(memory).rjust(8)}')","metadata":{"execution":{"iopub.status.busy":"2022-06-05T06:30:01.864837Z","iopub.execute_input":"2022-06-05T06:30:01.865205Z","iopub.status.idle":"2022-06-05T06:38:23.338403Z","shell.execute_reply.started":"2022-06-05T06:30:01.865171Z","shell.execute_reply":"2022-06-05T06:38:23.337193Z"},"trusted":true},"execution_count":null,"outputs":[]}]}