{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11228175,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# EDA of Sequence files\n\nThis notebook explores the contents of the sequecne files, starting from the the training one.\n\nThe format of this file is a bit complex to understand, because the last column contains multi-lines entries, in the FASTA format.\n\nFor example, in this entry, we have one sequence for Chain A of 1SCL_A:\n\n```\n1SCL_A,GGGUGCUCAGUACGAGAGGAACCGCACCC,1995-01-26,\"THE SARCIN-RICIN LOOP, A MODULAR RNA\",\">1SCL_1|Chain A|RNA SARCIN-RICIN LOOP|Rattus norvegicus (10116)\nGGGUGCUCAGUACGAGAGGAACCGCACCC\n\"\n```\n\nThis is another entry, with two sequences. \n\nThe first sequence is 1HMH_1, encoding Chains A, C, E of 1_HMH_E, with sequence GGCGACCCUGAUGAGGCCGAAAGGCCGAAACCGU. \n\nThe second sequence is 1HMH_2, and it encodes Chains B, D, F, and it reads ACGGTCGGTCGCC.\n\n```\n\n1HMH_E,GGCGACCCUGAUGAGGCCGAAAGGCCGAAACCGU,1995-12-07,THREE-DIMENSIONAL STRUCTURE OF A HAMMERHEAD RIBOZYME,\">1HMH_1|Chains A, C, E|HAMMERHEAD RIBOZYME-RNA STRAND|\nGGCGACCCUGAUGAGGCCGAAAGGCCGAAACCGU\n>1HMH_2|Chains B, D, F|HAMMERHEAD RIBOZYME-DNA STRAND|\nACGGTCGGTCGCC\n\"\n\n```\n\nFor this second example, according to the competition home page, we only need to make a prediction for the first chain, with sequence GGCGACCCUGAUGAGGCCGAAAGGCCGAAACCGU. We can ignore the second sequence, during the prediction, although it may be useful to use it for training.\n","metadata":{}},{"cell_type":"markdown","source":"## Original Description of the Sequence files\n\nThe following is the description of the sequences files, copy&pasted from the [competition home page](https://www.kaggle.com/competitions/stanford-rna-3d-folding/data)\n\n\\[train/validation/test]_sequences.csv - the target sequences of the RNA molecules.\n\n- target_id - (string) An arbitrary identifier. In train_sequences.csv, this is formatted as pdb_id_chain_id, where pdb_id is the id of the entry in the Protein Data Bank and chain_id is the chain id of the monomer in the pdb file.\n- sequence - (string) The RNA sequence. For test_sequences.csv, this is guaranteed to be a string of A, C, G, and U. For some train_sequences.csv, other characters may appear.\n- temporal_cutoff - (string) The date in yyyy-mm-dd format that the sequence was published. See Additional Notes.\n- description - (string) Details of the origins of the sequence. For a few targets, additional information on small molecule ligands bound to the RNA is included. You don't need to make predictions for these ligand coordinates.\n- all_sequences - (string) FASTA-formatted sequences of all molecular chains present in the experimentally solved structure. In a few cases this may include multiple copies of the target RNA (look for the word \"Chains\" in the header) and/or partners like other RNAs or proteins or DNA. You don't need to make predictions for all these molecules; if you do, just submit predictions for sequence. Some entries are blank.\n","metadata":{}},{"cell_type":"markdown","source":"### Contents of train_sequences.csv\n\nThese are the first lines of the train_sequences.file. \n\nThis is a CSV file, but the last column contains multi-line entries, in the FASTA format.\n\nNotice how some entries, like 1HMH_E, have multiple sequences.\n","metadata":{}},{"cell_type":"code","source":"!head -n 25 /kaggle/input/stanford-rna-3d-folding/train_sequences.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T11:42:23.423242Z","iopub.execute_input":"2025-03-05T11:42:23.423581Z","iopub.status.idle":"2025-03-05T11:42:23.549548Z","shell.execute_reply.started":"2025-03-05T11:42:23.423556Z","shell.execute_reply":"2025-03-05T11:42:23.548268Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Reading this file into a pandas dataframe\n\nWe need to use a multi-line approach to read all the entries in pandas.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport csv\n\n# Read only the first four columns, ignoring \"all_sequences\"\ndf = pd.read_csv(\n    \"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\",\n    engine=\"python\",\n    quoting=csv.QUOTE_MINIMAL,\n    usecols=[\"target_id\", \"sequence\", \"temporal_cutoff\", \"description\", \"all_sequences\"]\n)\n\nprint(df.head(10))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T11:46:56.160885Z","iopub.execute_input":"2025-03-05T11:46:56.161235Z","iopub.status.idle":"2025-03-05T11:46:56.214950Z","shell.execute_reply.started":"2025-03-05T11:46:56.161212Z","shell.execute_reply":"2025-03-05T11:46:56.214030Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.loc[df.target_id.str.contains(\"1HMH\")]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T11:45:06.191252Z","iopub.execute_input":"2025-03-05T11:45:06.191586Z","iopub.status.idle":"2025-03-05T11:45:06.202845Z","shell.execute_reply.started":"2025-03-05T11:45:06.191562Z","shell.execute_reply":"2025-03-05T11:45:06.201728Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# This should contain two entries for 1HMH_1 and 1HMH_2, but it seems the latter is lost. We do not need it for the prediction anyways.\ndf.loc[df.target_id.str.contains(\"1HMH\")].all_sequences","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T11:45:55.815482Z","iopub.execute_input":"2025-03-05T11:45:55.815852Z","iopub.status.idle":"2025-03-05T11:45:55.823895Z","shell.execute_reply.started":"2025-03-05T11:45:55.815815Z","shell.execute_reply":"2025-03-05T11:45:55.822849Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.loc[df.target_id.str.contains(\">\")]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T11:47:11.714682Z","iopub.execute_input":"2025-03-05T11:47:11.715050Z","iopub.status.idle":"2025-03-05T11:47:11.724446Z","shell.execute_reply.started":"2025-03-05T11:47:11.715021Z","shell.execute_reply":"2025-03-05T11:47:11.723423Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.head(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T11:45:34.678192Z","iopub.execute_input":"2025-03-05T11:45:34.678539Z","iopub.status.idle":"2025-03-05T11:45:34.689392Z","shell.execute_reply.started":"2025-03-05T11:45:34.678514Z","shell.execute_reply":"2025-03-05T11:45:34.688543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}}]}