{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":12993472,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport time\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"aa1356b4-81aa-4384-8303-38576393d6f4","_cell_guid":"75e0f525-7657-493c-9ff5-8f2662cf4468","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-15T16:12:35.659731Z","iopub.execute_input":"2025-07-15T16:12:35.661242Z","iopub.status.idle":"2025-07-15T16:12:35.671564Z","shell.execute_reply.started":"2025-07-15T16:12:35.661187Z","shell.execute_reply":"2025-07-15T16:12:35.670264Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = pd.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/train.parquet\")\ntest = pd.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/test.parquet\")","metadata":{"_uuid":"65a0ae85-0633-4167-bfa3-aed94aae1c34","_cell_guid":"300661e9-7820-4615-83b0-ff95afe82e16","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-15T16:12:35.674082Z","iopub.execute_input":"2025-07-15T16:12:35.675003Z","execution_failed":"2025-07-15T16:12:49.556Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Hi! I recently joined this competition and I know most of you are already past this stage of EDA, but I think this is interesting enough to share here. There are multiple columns that appear to be duplicates of each other. This means they do not contribute signal to our prediction and slow down training, and therefore have to be dropped\n\nThe problem at hand would now be the identification of duplicate columns. Initially, I went about this problem the direct way; comparing different columns and adding those that were duplicates of each other to a set. The column names in this set would then be used to drop duplicate columns from the dataset, as sets only store non-repeating values.","metadata":{"_uuid":"b7be5cba-45f0-4641-a160-c588af039f0c","_cell_guid":"0791ec13-a433-4cae-9eee-620faf834418","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"cols = train.columns\n\nstart_first = time.time()\ndef handle_duplicates(df):\n    dupes  = set()\n    \n    for i in range(len(cols)):\n        for j in range(i+1,len(cols)):\n            if (df.iloc[:,i] == df.iloc[:,j]).all():\n                dupes.add(cols[j])\n\n    return (list(dupes))\n\nduplicates = handle_duplicates(train)\nend_first = time.time()\n#print(f'Here is the list of duplicated columns: {duplicates}') \n#print(f'There are {len(duplicates)} columns that have direct duplicates')\nprint(f\"The first method took {end_first-start_first:.2f} seconds\")","metadata":{"_uuid":"c70cb5be-1ea6-4e6d-a35a-106ca1149ab6","_cell_guid":"147327f9-e6de-4e0b-abbd-d706bd8bb793","trusted":true,"collapsed":false,"execution":{"execution_failed":"2025-07-15T16:12:49.557Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The biggest issue with this approach is that it eats up a lot of compute and takes a long time to process due to the sheer size of the dataset. I therefore propose the use of hash values.\n\nThe individual elements of each column are hashed, and then a sum is computed and appended to a hash_sums list. There will be as many elements as there are columns in this list. The values in this list are then compared. The main assumption (we’ll see why it's an assumption later on in the notebook) is that if two columns have the same hash value, then the contents of those columns are the same and they are duplicates!\n\nSo rather than comparing columns across the entire dataset, we only compare their hash values; saving a lot of time.","metadata":{"_uuid":"5f7d965a-ce5e-47cc-b55e-43f09e48dfcb","_cell_guid":"ba4ae0f6-d193-439a-8666-63d35d058ac8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"#Hashing\nstart = time.time()\nhash_list = train.apply(lambda x :pd.util.hash_pandas_object(x,index=False).sum())\n\n#Extract the indices of the non repeating hashes\n_,hash_indices = np.unique(hash_list,return_index=True)\n\n#now our dataset doesn't contain any duplicates\ntrain = train.iloc[:,sorted(hash_indices)]\nend = time.time()\n\nprint(f\"The hashing method took {end-start:.2f} seconds\")","metadata":{"_uuid":"57980468-bfa1-4eb3-bf76-a6e29530bc33","_cell_guid":"58a90195-0107-49be-a73f-87497070ce96","trusted":true,"collapsed":false,"execution":{"execution_failed":"2025-07-15T16:12:49.557Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This yields the same result, for this case, as the first approach in a much shorter time!. We can prove this by comparing the number of columns after filtering out duplicates.","metadata":{"_uuid":"c41fd230-a43a-4e76-b546-0989f26fcdcb","_cell_guid":"7802e0db-d24a-49e9-9af6-2c2062ca8c95","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"#For the first method\nprint(f\"Unique columns after the first method: {len(cols)-len(duplicates)}\")\n\n#Second Method\nprint(f\"Unique columns after the second method: {len(hash_indices)}\")","metadata":{"_uuid":"04a6cdd9-9bc7-4bae-91a2-4103666d3fae","_cell_guid":"f1c876c2-f7aa-4ee8-a252-02b44c58b6e2","trusted":true,"collapsed":false,"execution":{"execution_failed":"2025-07-15T16:12:49.557Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The largest issue with this hashing approach is the risk of false positives from hash collisions. Since we're mapping from a high-dimensional space to a lower one, we are bound to have hashes that match without the columns actually being equal.\n\nI try to minimize this by first hashing every element in the column and then summing them up, rather than hashing the entire column at once. If you're still afraid of flagging columns as duplicates when they are not, then you can use this as a first-level filter; getting a list of columns that are probably duplicates. Then you can manually check them using the first method, but with less compute this time.\n\nIt’s my first time making one of these, and if you have any constructive criticism, please share it with me in the comments. Thank you for reading!!!!!","metadata":{"_uuid":"049d7b9f-87a5-406f-bc01-dae8fae290f3","_cell_guid":"a37a1816-e682-4b09-9a53-168aba8cee8a","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"","metadata":{"_uuid":"cc742038-bdef-49a5-be79-7f2e47e25762","_cell_guid":"03213910-8fae-4410-bc34-2393a3b60543","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}}]}