{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**I am not 100% sure**, but it seems that it is possible to reconstruct the order of the test rows.\nIn this notebook, I present the idea of sorting the test dataframe by date.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = pd.read_parquet('/kaggle/input/drw-crypto-market-prediction/train.parquet')\ntest_df = pd.read_parquet('/kaggle/input/drw-crypto-market-prediction/test.parquet')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:28:53.426773Z","iopub.execute_input":"2025-07-07T12:28:53.427195Z","iopub.status.idle":"2025-07-07T12:30:28.786393Z","shell.execute_reply.started":"2025-07-07T12:28:53.427172Z","shell.execute_reply":"2025-07-07T12:30:28.785258Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I paid attention to this [discussion post](https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/584485).\n\n> 2. Some columns are zero in train and not all zero in test.\n\nLet's identify such features.","metadata":{}},{"cell_type":"code","source":"train_zero_test_not_zero_cols = []\n\nfor col in train_df.columns:\n    if np.all(np.isclose(train_df[col].to_numpy(), 0.0)):\n        if not np.all(np.isclose(test_df[col].to_numpy(), 0.0)):\n            train_zero_test_not_zero_cols.append(col)\n\nprint(train_zero_test_not_zero_cols)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:30:28.789094Z","iopub.execute_input":"2025-07-07T12:30:28.789368Z","iopub.status.idle":"2025-07-07T12:30:32.391133Z","shell.execute_reply.started":"2025-07-07T12:30:28.789349Z","shell.execute_reply":"2025-07-07T12:30:32.390127Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in train_zero_test_not_zero_cols:\n    print(col, np.mean(np.isclose(test_df[col].to_numpy(), 0.0)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:30:32.392339Z","iopub.execute_input":"2025-07-07T12:30:32.392708Z","iopub.status.idle":"2025-07-07T12:30:32.427123Z","shell.execute_reply.started":"2025-07-07T12:30:32.392683Z","shell.execute_reply":"2025-07-07T12:30:32.426216Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It is quite natural to assume that, in every column, zeros should be before non-zeros.\nSo, the array","metadata":{}},{"cell_type":"code","source":"y_true = test_df[train_zero_test_not_zero_cols].astype(bool).sum(axis=1)\ny_true","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:30:32.427995Z","iopub.execute_input":"2025-07-07T12:30:32.428328Z","iopub.status.idle":"2025-07-07T12:30:32.524917Z","shell.execute_reply.started":"2025-07-07T12:30:32.428297Z","shell.execute_reply":"2025-07-07T12:30:32.524073Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"should be monotonically increasing.\n\nHow could the test rows have been shuffled? There is a chance that a simple shuffle with a specific *random_state* was applied. Let's check this hypothesis.","metadata":{}},{"cell_type":"code","source":"import tqdm\nfrom sklearn.metrics import accuracy_score\n\nseeds = np.arange(1000)\nscores = []\n\nfor seed in tqdm.tqdm(seeds):\n    y_pred = y_true.sort_values()\n    y_pred = y_pred.sample(n=len(y_pred), random_state=seed)\n    scores.append(accuracy_score(y_true, y_pred))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:30:32.525704Z","iopub.execute_input":"2025-07-07T12:30:32.525911Z","iopub.status.idle":"2025-07-07T12:31:59.231878Z","shell.execute_reply.started":"2025-07-07T12:30:32.525895Z","shell.execute_reply":"2025-07-07T12:31:59.230845Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(16, 4))\nplt.plot(seeds, scores)\nplt.title('random_state with max accuracy: ' + str(np.argmax(scores)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:31:59.233372Z","iopub.execute_input":"2025-07-07T12:31:59.234447Z","iopub.status.idle":"2025-07-07T12:31:59.697972Z","shell.execute_reply.started":"2025-07-07T12:31:59.234411Z","shell.execute_reply":"2025-07-07T12:31:59.696984Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can see a clear spike at ```random_state=700``` (also, the value ```700``` seems like it was chosen by a human).\nNow the question is: why isn't the accuracy exactly ```1```?\nIt turns out that our assumption (zeros should be before non-zeros) is not fully correct.\nWe can take a look at similar features in the training set.","metadata":{}},{"cell_type":"code","source":"zeroes_shares_df = train_df.astype(bool).mean(axis=0)\nzeroes_shares_df[(zeroes_shares_df > 0.4) & (zeroes_shares_df < 0.6)]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:31:59.699036Z","iopub.execute_input":"2025-07-07T12:31:59.699336Z","iopub.status.idle":"2025-07-07T12:32:02.257164Z","shell.execute_reply.started":"2025-07-07T12:31:59.699304Z","shell.execute_reply":"2025-07-07T12:32:02.256089Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df[['X526', 'X589', 'X778', 'X786']].iloc[:100000].plot(subplots=True, figsize=(16, 12))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:32:02.260494Z","iopub.execute_input":"2025-07-07T12:32:02.260830Z","iopub.status.idle":"2025-07-07T12:32:04.692951Z","shell.execute_reply.started":"2025-07-07T12:32:02.260809Z","shell.execute_reply":"2025-07-07T12:32:04.691931Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Indeed, the assumption that all zeros should precede non-zero values values is not 100% correct. That's why our accuracy was a bit lower than ```1```.","metadata":{}},{"cell_type":"markdown","source":"To sort the test dataframe, we can do the following:","metadata":{}},{"cell_type":"code","source":"def sort_test_df_by_time(df):\n    assert len(df.shape) == 2\n    assert df.shape[0] == 538150\n\n    n = df.shape[0]\n    t = pd.Series(np.arange(n))\n    t = t.sample(n=n, random_state=700)\n\n    t = pd.Series(np.arange(n), index=t.to_numpy()).sort_index()\n    return df.iloc[t.to_numpy()]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T13:06:40.965204Z","iopub.execute_input":"2025-07-07T13:06:40.965665Z","iopub.status.idle":"2025-07-07T13:06:40.972248Z","shell.execute_reply.started":"2025-07-07T13:06:40.965586Z","shell.execute_reply":"2025-07-07T13:06:40.971178Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sorted_test_df = sort_test_df_by_time(test_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:32:04.700918Z","iopub.execute_input":"2025-07-07T12:32:04.701229Z","iopub.status.idle":"2025-07-07T12:32:10.280684Z","shell.execute_reply.started":"2025-07-07T12:32:04.701205Z","shell.execute_reply":"2025-07-07T12:32:10.279794Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, let's save the sorted test dataframe to a file.\nYou can add this notebook to the inputs and read this sorted test dataframe from it.","metadata":{}},{"cell_type":"code","source":"sorted_test_df.to_parquet('sorted_test.parquet')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T12:32:10.281636Z","iopub.execute_input":"2025-07-07T12:32:10.281962Z","iopub.status.idle":"2025-07-07T12:32:46.033897Z","shell.execute_reply.started":"2025-07-07T12:32:10.281936Z","shell.execute_reply":"2025-07-07T12:32:46.032511Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}