{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11512973,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hello, in this notebook I show that:\n\n・Some sequences in test data are identical\n\n・Some sequences in test data have identical data in train data\n\n\nTest data will be updated on April 23rd upon start of Model training phase so this does not directly affect our final results. However, this insight will be a reminder for us to watch out for overfitting/score drop after phase shift.\nAlso, might be a bit helpful to delete identical sequence in train data, since current test data will be added to train data upon the start of Model training phase.\nHope this helps!\n\nI found this out while looking for common substrings between train and test dataset. Some sequences,if not identical,have long common substrings, which I'm planning to share in another notebook later.\n\nOne more thing: I am a kaggle biginner aspiring to be an expert. I have learned about deep learning, but have never competed in featured competition like this. \nAny discussion(And of course, upvote:)  ) is more than welcome! ","metadata":{"_kg_hide-input":false}},{"cell_type":"markdown","source":"【📚Quick Summary of result📚】\nBelow is the list of all sequences in test_sequences:\n\nR1107: No identical sequence\n\nR1108: No identical sequence\n\nR1116: No identical sequence\n\nR1117v2: No identical sequence\n\nR1126: Identical with 8TVZ_C\n\nR1128: Identical with 8BTZ_A\n\nR1136: No identical sequence\n\nR1138: Identical with 7PTL_B\n\nR1149: Identical with 8UYS_A\n\nR1156: Identical with 8UYJ_A\n\nR1189:Identical with 7YR7_A&R1190\n\nR1190:Identical with 7YR7_A&R1189","metadata":{}},{"cell_type":"markdown","source":"7 out of 12 current test sequences have counterpart in train sequence, and among these 7 sequences, 2 are identical.","metadata":{}},{"cell_type":"markdown","source":"【💡More thoughts💡】\n\n・Why, in some models, scores are low although we are predicting the identical data as train data?\n\n・What is the potential effect to scores when train and test have exactly the same data?\n\n・What is organizer's intention behind this?","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2025-03-22T17:12:07.734910Z","iopub.execute_input":"2025-03-22T17:12:07.735221Z","iopub.status.idle":"2025-03-22T17:12:07.742856Z","shell.execute_reply.started":"2025-03-22T17:12:07.735195Z","shell.execute_reply":"2025-03-22T17:12:07.741361Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\ntrain_sequences=pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\")\ntrain_labels=pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.csv\")\ntest_sequences = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/test_sequences.csv\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"First, lets' look at sequence of R1126 and 8TVZ_C for example:","metadata":{}},{"cell_type":"code","source":"print(\"sequence of R1126 in test_sequences:\")\ntest_sequences.iloc[4,1]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"sequence of 8TVZ_C in train_sequences:\")\ntrain_sequences.iloc[781,1]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"They look similar. Let's see if they are identical:","metadata":{}},{"cell_type":"code","source":"if test_sequences.iloc[4,1] == train_sequences.iloc[781,1]:\n    print(\"R1126 and 8TVZ_C are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Doing the same for the rest of the test data(data without identical sequence are not mentioned below):","metadata":{}},{"cell_type":"code","source":"if test_sequences.iloc[5,1] == train_sequences.iloc[722,1]:\n    print(\"R1128 and 8BTZ_A are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if test_sequences.iloc[7,1] == train_sequences.iloc[709,1]:\n    print(\"R1138 and 7PTL_B are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if test_sequences.iloc[8,1] == train_sequences.iloc[759,1]:\n    print(\"R1149 and 8UYS_A are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if test_sequences.iloc[9,1] == train_sequences.iloc[760,1]:\n    print(\"R1156 and 8UYJ_A are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if test_sequences.iloc[10,1] == train_sequences.iloc[739,1]:\n    print(\"R1189 and 7YR7_A are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if test_sequences.iloc[10,1] == train_sequences.iloc[739,1]:\n    print(\"R1189 and 7YR7_A are the same sequences\")","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}