{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Loading data\nWe have a lot of data in this competition, so it makes sense to use polars, a pandas alternative that is much faster and memory efficient, while multithreaded out of the box","metadata":{}},{"cell_type":"code","source":"import polars as pl\ntraining_data=pl.read_csv(\"/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv\") ","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:15:20.871942Z","iopub.execute_input":"2023-10-29T23:15:20.872467Z","iopub.status.idle":"2023-10-29T23:15:41.967901Z","shell.execute_reply.started":"2023-10-29T23:15:20.872433Z","shell.execute_reply":"2023-10-29T23:15:41.966108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get rid of duplicates\n\nFollowing dropping duplicates, sorting by \"sequence_id\", \"experiment_type\" also allows us to format the data in a way where every 2 rows we have 2A3/DMS of the same sequence. This way we can also reshape the data into Nx2 later on. You can also drop duplicates along with different criteria, such as signal to noise, although here I just show the simplest method. ","metadata":{}},{"cell_type":"code","source":"#drop duplicates based on \"sequence_id\", \"experiment_type\"\nprint(\"before dropping duplicates data shape is:\",training_data.shape)\ntraining_data=training_data.unique(subset=[\"sequence_id\", \"experiment_type\"]).sort([\"sequence_id\", \"experiment_type\"])\nprint(\"after dropping duplicates data shape is:\",training_data.shape)\ntraining_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:15:41.970056Z","iopub.execute_input":"2023-10-29T23:15:41.970444Z","iopub.status.idle":"2023-10-29T23:16:04.005257Z","shell.execute_reply.started":"2023-10-29T23:15:41.970415Z","shell.execute_reply":"2023-10-29T23:16:04.003141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Filter data based on signal to noise\nFilter based on the condition that SN_filter of both DMS/2A3 has to 1, matching how test set filtering is done. I will also save the filtered version of training data, which can be directly downloaded/used as notebook output","metadata":{}},{"cell_type":"code","source":"import numpy as np\nSN=training_data['SN_filter'].to_numpy().astype('int32').reshape(-1,2)\nSN=SN.min(-1)\nSN=np.repeat(SN,2)\nprint(\"before filtering data shape is:\",training_data.shape)\nfiltered_data=training_data.filter(SN==1)\nprint(\"after filtering data shape is:\",filtered_data.shape)\n\nfiltered_data=filtered_data.drop([\"reads\",\n                                  \"signal_to_noise\",\n                                  \"SN_filter\"])\n\nfiltered_data.write_csv('train_QUICK_START.csv') #554 MB\n","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:04.007589Z","iopub.execute_input":"2023-10-29T23:16:04.008155Z","iopub.status.idle":"2023-10-29T23:16:18.160659Z","shell.execute_reply.started":"2023-10-29T23:16:04.008109Z","shell.execute_reply":"2023-10-29T23:16:18.158648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del training_data","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:18.164973Z","iopub.execute_input":"2023-10-29T23:16:18.165567Z","iopub.status.idle":"2023-10-29T23:16:18.181861Z","shell.execute_reply.started":"2023-10-29T23:16:18.165538Z","shell.execute_reply":"2023-10-29T23:16:18.178381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filtered_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:18.184215Z","iopub.execute_input":"2023-10-29T23:16:18.184858Z","iopub.status.idle":"2023-10-29T23:16:18.211666Z","shell.execute_reply.started":"2023-10-29T23:16:18.184793Z","shell.execute_reply":"2023-10-29T23:16:18.210107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling by averaging reactivities\n\nLet's create a model by using avg reactivities for each experiment type as predictions. ","metadata":{}},{"cell_type":"code","source":"#get data\nlength=206\nlabel_names=[f\"reactivity_{i+1:04}\" for i in range(length)]\nlabels=filtered_data[label_names].to_numpy().astype('float32').reshape(-1,2,206).transpose(0,2,1)\nlabels=labels.clip(0,1)\nlabels.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:18.21485Z","iopub.execute_input":"2023-10-29T23:16:18.215376Z","iopub.status.idle":"2023-10-29T23:16:27.332798Z","shell.execute_reply.started":"2023-10-29T23:16:18.215329Z","shell.execute_reply":"2023-10-29T23:16:27.331605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"avg_DMS=np.nanmean(labels[:,:,1])\navg_2A3=np.nanmean(labels[:,:,0])","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:27.334389Z","iopub.execute_input":"2023-10-29T23:16:27.335421Z","iopub.status.idle":"2023-10-29T23:16:27.594889Z","shell.execute_reply.started":"2023-10-29T23:16:27.335389Z","shell.execute_reply":"2023-10-29T23:16:27.594016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del labels","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:27.597141Z","iopub.execute_input":"2023-10-29T23:16:27.598102Z","iopub.status.idle":"2023-10-29T23:16:27.606345Z","shell.execute_reply.started":"2023-10-29T23:16:27.598061Z","shell.execute_reply":"2023-10-29T23:16:27.605201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test data formatting\nDo inference with our avg values and make a submission","metadata":{}},{"cell_type":"markdown","source":"## First let's load the test and sample sub to see what's going on","metadata":{}},{"cell_type":"code","source":"test=pl.read_csv(\"/kaggle/input/stanford-ribonanza-rna-folding/test_sequences.csv\")\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:27.607944Z","iopub.execute_input":"2023-10-29T23:16:27.609117Z","iopub.status.idle":"2023-10-29T23:16:30.380638Z","shell.execute_reply.started":"2023-10-29T23:16:27.609079Z","shell.execute_reply":"2023-10-29T23:16:30.379499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub=pl.read_csv(\"/kaggle/input/stanford-ribonanza-rna-folding/sample_submission.csv\")\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:30.384519Z","iopub.execute_input":"2023-10-29T23:16:30.384932Z","iopub.status.idle":"2023-10-29T23:16:51.396351Z","shell.execute_reply.started":"2023-10-29T23:16:30.384893Z","shell.execute_reply":"2023-10-29T23:16:51.395233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Now put our predictions into sample sub and save","metadata":{}},{"cell_type":"code","source":"sample_sub=sample_sub.with_columns(pl.lit(avg_DMS).alias(\"reactivity_DMS_MaP\"),\n                                   pl.lit(avg_2A3).alias(\"reactivity_2A3_MaP\"),)\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:16:51.398005Z","iopub.execute_input":"2023-10-29T23:16:51.399618Z","iopub.status.idle":"2023-10-29T23:16:52.963185Z","shell.execute_reply.started":"2023-10-29T23:16:51.39951Z","shell.execute_reply":"2023-10-29T23:16:52.962082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.write_csv(\"submission.csv\",float_precision=2)","metadata":{"execution":{"iopub.status.busy":"2023-10-29T23:21:50.555771Z","iopub.execute_input":"2023-10-29T23:21:50.556091Z","iopub.status.idle":"2023-10-29T23:22:50.735958Z","shell.execute_reply.started":"2023-10-29T23:21:50.556069Z","shell.execute_reply":"2023-10-29T23:22:50.734993Z"},"trusted":true},"execution_count":null,"outputs":[]}]}