{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Files\n**train_data.csv** - the training data\\\n**test_sequences.csv** - the test set sequences, without any columns associated with the ground truth.\\\n**sample_submission.csv** - a sample submission file in the correct format\\","metadata":{}},{"cell_type":"code","source":"# train = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')\n# test = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/test_sequences.csv')\n# submission = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/sample_submission.csv')\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# These files takes a very long time to be loaded due to their sizes\n\n> train_data.csv - 2.37 GB\\\n> test_sequences.csv - 316.99 MB\\\n> sample_submission.csv - 3.67 GB","metadata":{}},{"cell_type":"markdown","source":"### Thus, following are the methods to load these large '.csv' files. For the demonstration purpose, I will be loading only train_data.csv file","metadata":{}},{"cell_type":"markdown","source":"### 1. Using pandas.read_csv(chunksize)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport time","metadata":{"execution":{"iopub.status.busy":"2023-09-08T12:00:54.994624Z","iopub.execute_input":"2023-09-08T12:00:54.995056Z","iopub.status.idle":"2023-09-08T12:00:55.000589Z","shell.execute_reply.started":"2023-09-08T12:00:54.995023Z","shell.execute_reply":"2023-09-08T12:00:54.999741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s1 = time.time()\ntrain1 = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv', chunksize=1000)\ndf = pd.concat(train1)\ne1 = time.time()\n\nprint(\"Time using chunksize\", (e1-s1), \"sec\") \n\ndf.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-09-08T12:00:58.922399Z","iopub.execute_input":"2023-09-08T12:00:58.922833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. Using Dask","metadata":{}},{"cell_type":"code","source":"!pip install dask","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from dask import dataframe as dd\n  \ns2 = time.time()\ntrain2 = dd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')\ne2 = time.time()\n  \nprint(\"Time using Dask: \", (e2-s2), \"sec\")\n\ntrain2.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-09-08T11:09:58.918473Z","iopub.execute_input":"2023-09-08T11:09:58.918945Z","iopub.status.idle":"2023-09-08T11:10:03.363086Z","shell.execute_reply.started":"2023-09-08T11:09:58.918909Z","shell.execute_reply":"2023-09-08T11:10:03.361879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Columns\n*  **id** - integer (0,1,…) that identifies each sequence position in the sample submission.\\\n*  **id_min, id_max** - (integer) minimum and maximum id values for each test sequence.\\\n*  **sequence_id** - (string) An arbitrary identifier like 8cdfeef00 for each sequence.\\\n*  **sequence** - (string) Describes the RNA sequence, a combination of A, G, U, and C for each sample. Should be 115 to 457 characters long).\\\n*  **experiment_type** - (string) Either DMS_MaP or 2A3_MaP to describe the type of chemical mapping experiment that was used to generate each profile. References: DMS, 2A3.\\\n*  **dataset_name** - (string) name of high throughput sequencing dataset from which the reactivity profile was extracted.\\\n*  **reads** - (integer) Number of reads in the high throughput sequencing experiment that were assigned to the RNA sequence, and whose mutations were tabulated to compile the reactivity profile.\\\n*  **signal_to_noise** - (float) Signal/noise value for the profile, defined as mean( measurement value over probed nts )/mean( statistical error in measurement value over probed nts).\\\n*  **SN_filter** - (Boolean) 0 or 1 depending on whether the profile has signal_to_noise>1.0 and reads>100. For evaluation, only sequences whose DMS_MaP and 2A3_MaP profiles both pass this filter will be used to score submissions.\\\n*  **reactivity_0001, reactivity_0002,…** - (float) An array of floating point numbers of the train data, should have the same length as the RNA sequence, which defines the reactivity profile for the RNA. For sequences shorter than the maximum RNA length, positions that go beyond the sequence length have null. Several positions early and late in the sequence also cannot be probed due to technical reasons, and their reactivity values are null.\\\n*  **reactivity_error_0001, reactivity_error_0002,…** - (float) An array of floating point numbers, should have the same length as the corresponding reactivity_* columns, calculated errors in experimental values obtained in reactivity derived from counting statistics in the high-throughput sequencing experiment.\\\n*  **reactivity_DMS_MaP, reactivity_2A3_MaP** - (float) sample submission values.\\\n*  **future** - (Boolean) sequences whose data will be collected after the start of the competition (but before final scoring) are labeled as 1.","metadata":{}},{"cell_type":"markdown","source":"Reference: https://www.geeksforgeeks.org/working-with-large-csv-files-in-python/","metadata":{}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}