{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59575,"databundleVersionId":8060720,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Update\nI created a index as described below it can be found at this link: https://www.kaggle.com/datasets/devinanzelmo/uspto-explainable-ai-validation-index/","metadata":{"_uuid":"e8163c4a-2a8a-4d52-abda-5913ef5eafc7","_cell_guid":"128feff4-7777-4ebe-969c-e531e3d84cf5","jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"## Issue with the design of train_index\n\nThe train_index is not very helpful for doing validation. The issue is that for any given publication_number it does not contain enough of the neighbors to figure out if your query is working or not. As shown below at maximum 10 neighbors are included in the train_index for a publication_number. \n\nOne fix for the train_index is as follows\n* Pick 4000 publication_numbers from after 1975 (obviously not in test set)\n* Include all the neighbors for the above 4000 publication_numbers in the index\n* Include the list of 4000 publication_numbers in the dataset with the index\n\nThis will result in a 200k index like we have now but it will be much more useful for validation as we can see how many of the neighbors we are able to recover with a query. This will make working on the competition much easier (currently we would be guessing if our query is good).\n\n","metadata":{}},{"cell_type":"code","source":"import polars as pl\nimport json","metadata":{"execution":{"iopub.status.busy":"2024-04-26T11:51:59.212113Z","iopub.execute_input":"2024-04-26T11:51:59.212603Z","iopub.status.idle":"2024-04-26T11:51:59.607011Z","shell.execute_reply.started":"2024-04-26T11:51:59.212559Z","shell.execute_reply":"2024-04-26T11:51:59.605194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read the json file containing the publication_numbers contained in the train_index\nwith open(\"/kaggle/input/uspto-explainable-ai/train_index_patent_ids.json\", \"r\") as f:\n    train_index_ids = json.load(f)\ntrain_index_ids_dict = dict(zip(train_index_ids, [0]*len(train_index_ids)))","metadata":{"execution":{"iopub.status.busy":"2024-04-26T11:51:59.617307Z","iopub.execute_input":"2024-04-26T11:51:59.618202Z","iopub.status.idle":"2024-04-26T11:51:59.793093Z","shell.execute_reply.started":"2024-04-26T11:51:59.618127Z","shell.execute_reply":"2024-04-26T11:51:59.791825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get nearest neighbor rows corrosponding to the publication numbers in train_index\nnearest_neighbors = pl.read_csv(\"/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv\")\n\n#nearest_neighbors_small = nearest_neighbors.filter(pl.col(\"publication_number\").is_in(train_index_ids))\n#del nearest_neighbors","metadata":{"execution":{"iopub.status.busy":"2024-04-26T11:51:59.975738Z","iopub.execute_input":"2024-04-26T11:51:59.977336Z","iopub.status.idle":"2024-04-26T11:52:54.363985Z","shell.execute_reply.started":"2024-04-26T11:51:59.977283Z","shell.execute_reply":"2024-04-26T11:52:54.362195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# count the number of neighbors in train_index for each publication_number\ncounts = list()\nfor i in range(nearest_neighbors.shape[0]):\n    count = 0\n    row = nearest_neighbors[i,:].transpose()[:,0].to_list()\n    for j in row[1:]:\n        if j in train_index_ids_dict:\n            count += 1\n    counts.append([row[0], count])","metadata":{"execution":{"iopub.status.busy":"2024-04-26T11:56:12.196899Z","iopub.execute_input":"2024-04-26T11:56:12.197445Z","iopub.status.idle":"2024-04-26T12:25:16.356537Z","shell.execute_reply.started":"2024-04-26T11:56:12.197404Z","shell.execute_reply":"2024-04-26T12:25:16.354328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"counts_df = pl.DataFrame(counts, schema=[\"publication_number\", \"count\"])","metadata":{"execution":{"iopub.status.busy":"2024-04-26T12:25:16.360537Z","iopub.execute_input":"2024-04-26T12:25:16.361133Z","iopub.status.idle":"2024-04-26T12:25:31.704355Z","shell.execute_reply.started":"2024-04-26T12:25:16.361088Z","shell.execute_reply":"2024-04-26T12:25:31.702355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sort and display some counts\ncounts_df = counts_df.sort(by=\"count\", descending=True)\ncounts_df","metadata":{"execution":{"iopub.status.busy":"2024-04-26T12:25:31.706953Z","iopub.execute_input":"2024-04-26T12:25:31.707627Z","iopub.status.idle":"2024-04-26T12:25:33.001724Z","shell.execute_reply.started":"2024-04-26T12:25:31.707567Z","shell.execute_reply":"2024-04-26T12:25:33.000102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}