{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59575,"databundleVersionId":8060720,"sourceType":"competition"},{"sourceId":8246447,"sourceType":"datasetVersion","datasetId":4892374},{"sourceId":8323913,"sourceType":"datasetVersion","datasetId":4944579},{"sourceId":8413600,"sourceType":"datasetVersion","datasetId":5007812},{"sourceId":8479599,"sourceType":"datasetVersion","datasetId":4517815},{"sourceId":174185912,"sourceType":"kernelVersion"},{"sourceId":27825,"sourceType":"modelInstanceVersion","modelInstanceId":22009}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":309.540121,"end_time":"2024-05-07T13:59:54.129168","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-05-07T13:54:44.589047","version":"2.5.0"},"widgets":{"application/vnd.jupyter.widget-state+json":{"state":{"113ad24a0a824d4db61bdb6f41ad1926":{"model_module":"@jupyter-widgets/base","model_module_version":"1.2.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"1.2.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"overflow_x":null,"overflow_y":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2e051952b1ef4e2ea662a8415068011a":{"model_module":"@jupyter-widgets/base","model_module_version":"1.2.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"1.2.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"overflow_x":null,"overflow_y":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"4692e1a15c4f4c88ad61364cb5dc550d":{"model_module":"@jupyter-widgets/base","model_module_version":"1.2.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"1.2.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"overflow_x":null,"overflow_y":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"52dc50d35a9c40f0bee442a033d1f2cc":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"DescriptionStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"DescriptionStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"StyleView","description_width":""}},"5a1b7fb9f35a4462a10833cff7f1f167":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"1.5.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_d3e1316f84004457975ddbb19ada7946","IPY_MODEL_84a61eb3b7b2467bae87d65efdac7e16","IPY_MODEL_bc53c1da3adf4a23a53f59562650b161"],"layout":"IPY_MODEL_2e051952b1ef4e2ea662a8415068011a"}},"84a61eb3b7b2467bae87d65efdac7e16":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"1.5.0","_view_name":"ProgressView","bar_style":"success","description":"","description_tooltip":null,"layout":"IPY_MODEL_a426257544b2424a9ccd0001fa5add5c","max":2,"min":0,"orientation":"horizontal","style":"IPY_MODEL_8c9ba10a08c94efcb8eb307a2f027d07","value":2}},"8570018c7fe24b3594ec85ce309ddab2":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"DescriptionStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"DescriptionStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"StyleView","description_width":""}},"8c9ba10a08c94efcb8eb307a2f027d07":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"a426257544b2424a9ccd0001fa5add5c":{"model_module":"@jupyter-widgets/base","model_module_version":"1.2.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"1.2.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"1.2.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"overflow_x":null,"overflow_y":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"bc53c1da3adf4a23a53f59562650b161":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"1.5.0","_view_name":"HTMLView","description":"","description_tooltip":null,"layout":"IPY_MODEL_4692e1a15c4f4c88ad61364cb5dc550d","placeholder":"​","style":"IPY_MODEL_52dc50d35a9c40f0bee442a033d1f2cc","value":" 2/2 [01:35&lt;00:00, 44.31s/it]"}},"d3e1316f84004457975ddbb19ada7946":{"model_module":"@jupyter-widgets/controls","model_module_version":"1.5.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"1.5.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"1.5.0","_view_name":"HTMLView","description":"","description_tooltip":null,"layout":"IPY_MODEL_113ad24a0a824d4db61bdb6f41ad1926","placeholder":"​","style":"IPY_MODEL_8570018c7fe24b3594ec85ce309ddab2","value":"Loading checkpoint shards: 100%"}}},"version_major":2,"version_minor":0}}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=\"https://keras.io/img/logo-small.png\" alt=\"Keras logo\" width=\"100\"><br/>\nThis starter notebook is provided by the Keras team.</center>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# USPTO - Explainable AI for Patent Professionals with [KerasNLP](https://github.com/keras-team/keras-nlp) and [Keras](https://github.com/keras-team/keras)\n\n<div align=\"center\">\n    <img src=\"https://upload.wikimedia.org/wikipedia/commons/7/75/Seal_of_the_United_States_Patent_and_Trademark_Office.svg\" width=\"150\">\n</div>\n\nIn this competition, our aim is to enable patent professionals to interpret the results of AI-powered searches in familiar language and syntax through Boolean search queries. These queries, which effectively characterize groups of patents, are well known to patent professionals. Specifically, we need to create a model that will generate binary search queries from a patent and its neighbor $50$ patents, which will return the same set of $50$ neighbor patents when searched in the USPTO database. The problem in this competition can be visually described as below:\n\n<div align=\"center\"><img src=\"https://i.postimg.cc/q7jYCHLY/USPTO-problem.png\" width=\"550\"></div>\n\nIn this notebook, we will demonstrate how to leverage Large Language Models (LLMs) to approach this problem with an iterative comparison method to extract effective queries. Specifically, this notebook uses the **Gemma 1.1 2B** instruction tuned model to perform query extraction from the patents.\n\n**Did you know**: This notebook is backend-agnostic, which means it supports TensorFlow, PyTorch, and JAX backends. However, the best performance can be achieved with `JAX`. KerasNLP and Keras enable the choice of the preferred backend. Explore further details on [Keras](https://keras.io/keras_3/).\n\n**Note**: For a deeper understanding of KerasNLP, refer to the [KerasNLP guides](https://keras.io/keras_nlp/).\n","metadata":{}},{"cell_type":"markdown","source":"# 🛠️ | Install Libraries\n\nWe need to install `whoosh` and `whoosh_utils` libraries which will be used for processing our model output. As we don't have access to internet during inference, we will be installing this library from our local files.\n","metadata":{}},{"cell_type":"code","source":"# Install whoosh and whoosh_utils library\n!pip install /kaggle/input/uspto-whoosh-reloaded-2-7-5-patched/Whoosh_Reloaded-2.7.5-py2.py3-none-any.whl\n!sed 's:/kaggle/input/whoosh-wheel-2-7-4/Whoosh-2.7.4-py2.py3-none-any.whl:whoosh-reloaded==2.7.5:g' /kaggle/usr/lib/whoosh_utils/whoosh_utils.py > whoosh_utils.py","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-06-07T05:01:28.602017Z","iopub.execute_input":"2024-06-07T05:01:28.602575Z","iopub.status.idle":"2024-06-07T05:01:42.318729Z","shell.execute_reply.started":"2024-06-07T05:01:28.602539Z","shell.execute_reply":"2024-06-07T05:01:42.317529Z"},"trusted":true},"execution_count":3,"outputs":[{"name":"stderr","text":"/opt/conda/lib/python3.10/pty.py:89: RuntimeWarning: os.fork() was called. os.fork() is incompatible with multithreaded code, and JAX is multithreaded, so this will likely lead to a deadlock.\n  pid, fd = os.forkpty()\n","output_type":"stream"},{"name":"stdout","text":"Processing /kaggle/input/uspto-whoosh-reloaded-2-7-5-patched/Whoosh_Reloaded-2.7.5-py2.py3-none-any.whl\nRequirement already satisfied: cached-property in /opt/conda/lib/python3.10/site-packages (from Whoosh-Reloaded==2.7.5) (1.5.2)\nWhoosh-Reloaded is already installed with the same version as the provided wheel. Use --force-reinstall to force an installation of the wheel.\n","output_type":"stream"}]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install /kaggle/input/uspto-whoosh-reloaded-2-7-5-patched/Whoosh_Reloaded-2.7.5-py2.py3-none-any.whl\n!sed 's:/kaggle/input/whoosh-wheel-2-7-4/Whoosh-2.7.4-py2.py3-none-any.whl:whoosh-reloaded==2.7.5:g' /kaggle/usr/lib/whoosh_utils/whoosh_utils.py > whoosh_utils.py","metadata":{"execution":{"iopub.status.busy":"2024-06-07T04:44:37.028243Z","iopub.execute_input":"2024-06-07T04:44:37.02908Z","iopub.status.idle":"2024-06-07T04:44:51.514597Z","shell.execute_reply.started":"2024-06-07T04:44:37.029037Z","shell.execute_reply":"2024-06-07T04:44:51.513391Z"},"trusted":true},"execution_count":1,"outputs":[{"name":"stdout","text":"Processing /kaggle/input/uspto-whoosh-reloaded-2-7-5-patched/Whoosh_Reloaded-2.7.5-py2.py3-none-any.whl\nRequirement already satisfied: cached-property in /opt/conda/lib/python3.10/site-packages (from Whoosh-Reloaded==2.7.5) (1.5.2)\nInstalling collected packages: Whoosh-Reloaded\nSuccessfully installed Whoosh-Reloaded-2.7.5\n","output_type":"stream"}]},{"cell_type":"markdown","source":"# 📚 | Import Libraries ","metadata":{}},{"cell_type":"code","source":"import os\nos.environ[\"KERAS_BACKEND\"] = \"jax\"  # or \"tensorflow\" or \"torch\"\n\nimport keras_nlp\nimport keras\n\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport gc\n\nimport re\nimport whoosh_utils\nimport whoosh","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-06-07T05:00:55.983548Z","iopub.execute_input":"2024-06-07T05:00:55.984236Z","iopub.status.idle":"2024-06-07T05:01:19.913169Z","shell.execute_reply.started":"2024-06-07T05:00:55.984202Z","shell.execute_reply":"2024-06-07T05:01:19.911745Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":2,"outputs":[{"name":"stderr","text":"2024-06-07 05:00:57.718556: E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:9261] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered\n2024-06-07 05:00:57.718711: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:607] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered\n2024-06-07 05:00:57.847660: E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1515] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered\n","output_type":"stream"},{"name":"stdout","text":"Requirement already satisfied: whoosh-reloaded==2.7.5 in /opt/conda/lib/python3.10/site-packages (2.7.5)\nRequirement already satisfied: cached-property in /opt/conda/lib/python3.10/site-packages (from whoosh-reloaded==2.7.5) (1.5.2)\n","output_type":"stream"},{"name":"stderr","text":"\u001b[31mERROR: Operation cancelled by user\u001b[0m\u001b[31m\n\u001b[0m","output_type":"stream"},{"traceback":["\u001b[0;31m---------------------------------------------------------------------------\u001b[0m","\u001b[0;31mKeyboardInterrupt\u001b[0m                         Traceback (most recent call last)","Cell \u001b[0;32mIn[2], line 13\u001b[0m\n\u001b[1;32m     10\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mgc\u001b[39;00m\n\u001b[1;32m     12\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mre\u001b[39;00m\n\u001b[0;32m---> 13\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mwhoosh_utils\u001b[39;00m\n\u001b[1;32m     14\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mwhoosh\u001b[39;00m\n","File \u001b[0;32m/kaggle/working/whoosh_utils.py:25\u001b[0m\n\u001b[1;32m     22\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01msubprocess\u001b[39;00m\n\u001b[1;32m     24\u001b[0m \u001b[38;5;28;01mif\u001b[39;00m os\u001b[38;5;241m.\u001b[39mpath\u001b[38;5;241m.\u001b[39mexists(\u001b[38;5;124m'\u001b[39m\u001b[38;5;124m/kaggle/working\u001b[39m\u001b[38;5;124m'\u001b[39m):\n\u001b[0;32m---> 25\u001b[0m     \u001b[43msubprocess\u001b[49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43mrun\u001b[49m\u001b[43m(\u001b[49m\u001b[38;5;124;43m'\u001b[39;49m\u001b[38;5;124;43mpip install whoosh-reloaded==2.7.5\u001b[39;49m\u001b[38;5;124;43m'\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[43mshell\u001b[49m\u001b[38;5;241;43m=\u001b[39;49m\u001b[38;5;28;43;01mTrue\u001b[39;49;00m\u001b[43m)\u001b[49m\n\u001b[1;32m     27\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mwhoosh\u001b[39;00m\u001b[38;5;21;01m.\u001b[39;00m\u001b[38;5;21;01manalysis\u001b[39;00m\n\u001b[1;32m     28\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mwhoosh\u001b[39;00m\u001b[38;5;21;01m.\u001b[39;00m\u001b[38;5;21;01mcollectors\u001b[39;00m\n","File \u001b[0;32m/opt/conda/lib/python3.10/subprocess.py:505\u001b[0m, in \u001b[0;36mrun\u001b[0;34m(input, capture_output, timeout, check, *popenargs, **kwargs)\u001b[0m\n\u001b[1;32m    503\u001b[0m \u001b[38;5;28;01mwith\u001b[39;00m Popen(\u001b[38;5;241m*\u001b[39mpopenargs, \u001b[38;5;241m*\u001b[39m\u001b[38;5;241m*\u001b[39mkwargs) \u001b[38;5;28;01mas\u001b[39;00m process:\n\u001b[1;32m    504\u001b[0m     \u001b[38;5;28;01mtry\u001b[39;00m:\n\u001b[0;32m--> 505\u001b[0m         stdout, stderr \u001b[38;5;241m=\u001b[39m \u001b[43mprocess\u001b[49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43mcommunicate\u001b[49m\u001b[43m(\u001b[49m\u001b[38;5;28;43minput\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[43mtimeout\u001b[49m\u001b[38;5;241;43m=\u001b[39;49m\u001b[43mtimeout\u001b[49m\u001b[43m)\u001b[49m\n\u001b[1;32m    506\u001b[0m     \u001b[38;5;28;01mexcept\u001b[39;00m TimeoutExpired \u001b[38;5;28;01mas\u001b[39;00m exc:\n\u001b[1;32m    507\u001b[0m         process\u001b[38;5;241m.\u001b[39mkill()\n","File \u001b[0;32m/opt/conda/lib/python3.10/subprocess.py:1146\u001b[0m, in \u001b[0;36mPopen.communicate\u001b[0;34m(self, input, timeout)\u001b[0m\n\u001b[1;32m   1144\u001b[0m         stderr \u001b[38;5;241m=\u001b[39m \u001b[38;5;28mself\u001b[39m\u001b[38;5;241m.\u001b[39mstderr\u001b[38;5;241m.\u001b[39mread()\n\u001b[1;32m   1145\u001b[0m         \u001b[38;5;28mself\u001b[39m\u001b[38;5;241m.\u001b[39mstderr\u001b[38;5;241m.\u001b[39mclose()\n\u001b[0;32m-> 1146\u001b[0m     \u001b[38;5;28;43mself\u001b[39;49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43mwait\u001b[49m\u001b[43m(\u001b[49m\u001b[43m)\u001b[49m\n\u001b[1;32m   1147\u001b[0m \u001b[38;5;28;01melse\u001b[39;00m:\n\u001b[1;32m   1148\u001b[0m     \u001b[38;5;28;01mif\u001b[39;00m timeout \u001b[38;5;129;01mis\u001b[39;00m \u001b[38;5;129;01mnot\u001b[39;00m \u001b[38;5;28;01mNone\u001b[39;00m:\n","File \u001b[0;32m/opt/conda/lib/python3.10/subprocess.py:1209\u001b[0m, in \u001b[0;36mPopen.wait\u001b[0;34m(self, timeout)\u001b[0m\n\u001b[1;32m   1207\u001b[0m     endtime \u001b[38;5;241m=\u001b[39m _time() \u001b[38;5;241m+\u001b[39m timeout\n\u001b[1;32m   1208\u001b[0m \u001b[38;5;28;01mtry\u001b[39;00m:\n\u001b[0;32m-> 1209\u001b[0m     \u001b[38;5;28;01mreturn\u001b[39;00m \u001b[38;5;28;43mself\u001b[39;49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43m_wait\u001b[49m\u001b[43m(\u001b[49m\u001b[43mtimeout\u001b[49m\u001b[38;5;241;43m=\u001b[39;49m\u001b[43mtimeout\u001b[49m\u001b[43m)\u001b[49m\n\u001b[1;32m   1210\u001b[0m \u001b[38;5;28;01mexcept\u001b[39;00m \u001b[38;5;167;01mKeyboardInterrupt\u001b[39;00m:\n\u001b[1;32m   1211\u001b[0m     \u001b[38;5;66;03m# https://bugs.python.org/issue25942\u001b[39;00m\n\u001b[1;32m   1212\u001b[0m     \u001b[38;5;66;03m# The first keyboard interrupt waits briefly for the child to\u001b[39;00m\n\u001b[1;32m   1213\u001b[0m     \u001b[38;5;66;03m# exit under the common assumption that it also received the ^C\u001b[39;00m\n\u001b[1;32m   1214\u001b[0m     \u001b[38;5;66;03m# generated SIGINT and will exit rapidly.\u001b[39;00m\n\u001b[1;32m   1215\u001b[0m     \u001b[38;5;28;01mif\u001b[39;00m timeout \u001b[38;5;129;01mis\u001b[39;00m \u001b[38;5;129;01mnot\u001b[39;00m \u001b[38;5;28;01mNone\u001b[39;00m:\n","File \u001b[0;32m/opt/conda/lib/python3.10/subprocess.py:1959\u001b[0m, in \u001b[0;36mPopen._wait\u001b[0;34m(self, timeout)\u001b[0m\n\u001b[1;32m   1957\u001b[0m \u001b[38;5;28;01mif\u001b[39;00m \u001b[38;5;28mself\u001b[39m\u001b[38;5;241m.\u001b[39mreturncode \u001b[38;5;129;01mis\u001b[39;00m \u001b[38;5;129;01mnot\u001b[39;00m \u001b[38;5;28;01mNone\u001b[39;00m:\n\u001b[1;32m   1958\u001b[0m     \u001b[38;5;28;01mbreak\u001b[39;00m  \u001b[38;5;66;03m# Another thread waited.\u001b[39;00m\n\u001b[0;32m-> 1959\u001b[0m (pid, sts) \u001b[38;5;241m=\u001b[39m \u001b[38;5;28;43mself\u001b[39;49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43m_try_wait\u001b[49m\u001b[43m(\u001b[49m\u001b[38;5;241;43m0\u001b[39;49m\u001b[43m)\u001b[49m\n\u001b[1;32m   1960\u001b[0m \u001b[38;5;66;03m# Check the pid and loop as waitpid has been known to\u001b[39;00m\n\u001b[1;32m   1961\u001b[0m \u001b[38;5;66;03m# return 0 even without WNOHANG in odd situations.\u001b[39;00m\n\u001b[1;32m   1962\u001b[0m \u001b[38;5;66;03m# http://bugs.python.org/issue14396.\u001b[39;00m\n\u001b[1;32m   1963\u001b[0m \u001b[38;5;28;01mif\u001b[39;00m pid \u001b[38;5;241m==\u001b[39m \u001b[38;5;28mself\u001b[39m\u001b[38;5;241m.\u001b[39mpid:\n","File \u001b[0;32m/opt/conda/lib/python3.10/subprocess.py:1917\u001b[0m, in \u001b[0;36mPopen._try_wait\u001b[0;34m(self, wait_flags)\u001b[0m\n\u001b[1;32m   1915\u001b[0m \u001b[38;5;250m\u001b[39m\u001b[38;5;124;03m\"\"\"All callers to this function MUST hold self._waitpid_lock.\"\"\"\u001b[39;00m\n\u001b[1;32m   1916\u001b[0m \u001b[38;5;28;01mtry\u001b[39;00m:\n\u001b[0;32m-> 1917\u001b[0m     (pid, sts) \u001b[38;5;241m=\u001b[39m \u001b[43mos\u001b[49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43mwaitpid\u001b[49m\u001b[43m(\u001b[49m\u001b[38;5;28;43mself\u001b[39;49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43mpid\u001b[49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[43mwait_flags\u001b[49m\u001b[43m)\u001b[49m\n\u001b[1;32m   1918\u001b[0m \u001b[38;5;28;01mexcept\u001b[39;00m \u001b[38;5;167;01mChildProcessError\u001b[39;00m:\n\u001b[1;32m   1919\u001b[0m     \u001b[38;5;66;03m# This happens if SIGCLD is set to be ignored or waiting\u001b[39;00m\n\u001b[1;32m   1920\u001b[0m     \u001b[38;5;66;03m# for child processes has otherwise been disabled for our\u001b[39;00m\n\u001b[1;32m   1921\u001b[0m     \u001b[38;5;66;03m# process.  This child is dead, we can't get the status.\u001b[39;00m\n\u001b[1;32m   1922\u001b[0m     pid \u001b[38;5;241m=\u001b[39m \u001b[38;5;28mself\u001b[39m\u001b[38;5;241m.\u001b[39mpid\n","\u001b[0;31mKeyboardInterrupt\u001b[0m: "],"ename":"KeyboardInterrupt","evalue":"","output_type":"error"}]},{"cell_type":"markdown","source":"## Library Version","metadata":{}},{"cell_type":"code","source":"print(\"Keras:\", keras.__version__)\nprint(\"KerasNLP:\", keras_nlp.__version__)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-30T08:45:29.169299Z","iopub.execute_input":"2024-05-30T08:45:29.169668Z","iopub.status.idle":"2024-05-30T08:45:29.175497Z","shell.execute_reply.started":"2024-05-30T08:45:29.169629Z","shell.execute_reply":"2024-05-30T08:45:29.174275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Keras:\", keras.__version__)\nprint(\"KerasNLP:\", keras_nlp.__version__)","metadata":{"execution":{"iopub.status.busy":"2024-06-07T04:50:58.725649Z","iopub.execute_input":"2024-06-07T04:50:58.726698Z","iopub.status.idle":"2024-06-07T04:50:58.731578Z","shell.execute_reply.started":"2024-06-07T04:50:58.726664Z","shell.execute_reply":"2024-06-07T04:50:58.730438Z"},"trusted":true},"execution_count":3,"outputs":[{"name":"stdout","text":"Keras: 3.3.3\nKerasNLP: 0.12.1\n","output_type":"stream"}]},{"cell_type":"markdown","source":"# ⚙️ | Configuration","metadata":{}},{"cell_type":"code","source":"class CFG:\n    seed = 42\n    dataset_path = \"/kaggle/input/ai-mathematical-olympiad-prize\"\n    preset = \"gemma_1.1_instruct_2b_en\" # name of pretrained Gemma\n    input_length = 1024 # max size of input sequence for training\n    output_length = 1200 # max size of output sequence\n    num_neighbors = 2 # how many neighbour patents to consider","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:45:29.177286Z","iopub.execute_input":"2024-05-30T08:45:29.177558Z","iopub.status.idle":"2024-05-30T08:45:29.187057Z","shell.execute_reply.started":"2024-05-30T08:45:29.177535Z","shell.execute_reply":"2024-05-30T08:45:29.186223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class CFG:\n    seed = 42\n    dataset_path = \"/kaggle/input/ai-mathematical-olympiad-prize\"\n    preset = \"gemma_1.1_instruct_2b_en\" # name of pretrained Gemma\n    input_length = 1024 # max size of input sequence for training\n    output_length = 1200 # max size of output sequence\n    num_neighbors = 2 # how many neighbour patents to consider","metadata":{"execution":{"iopub.status.busy":"2024-06-07T04:51:15.32414Z","iopub.execute_input":"2024-06-07T04:51:15.324785Z","iopub.status.idle":"2024-06-07T04:51:15.329533Z","shell.execute_reply.started":"2024-06-07T04:51:15.324753Z","shell.execute_reply":"2024-06-07T04:51:15.32848Z"},"trusted":true},"execution_count":6,"outputs":[]},{"cell_type":"markdown","source":"# ♻️ | Reproducibility \n\nSets value for random seed to produce similar result in each run.","metadata":{}},{"cell_type":"code","source":"keras.utils.set_random_seed(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:45:29.188204Z","iopub.execute_input":"2024-05-30T08:45:29.188494Z","iopub.status.idle":"2024-05-30T08:45:29.198914Z","shell.execute_reply.started":"2024-05-30T08:45:29.18846Z","shell.execute_reply":"2024-05-30T08:45:29.198097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"keras.utils.set_random_seed(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2024-06-07T04:53:00.625254Z","iopub.execute_input":"2024-06-07T04:53:00.626099Z","iopub.status.idle":"2024-06-07T04:53:00.630098Z","shell.execute_reply.started":"2024-06-07T04:53:00.626067Z","shell.execute_reply":"2024-06-07T04:53:00.629121Z"},"trusted":true},"execution_count":14,"outputs":[]},{"cell_type":"markdown","source":"# 📖 | Metadata\n\nThe competition dataset provides different information about patents and their neighboring patents. In this notebook, we will use the `Title` and the `Abstract` of the patents to extract search queries.\n\n## Files\n\n### **patent_metadata.parquet** → Metadata for the most recent patents\n- `publication_number` - The patent identifier.\n- `publication_date` - The date the patent was published.\n- `filing_date` - The date the patent was filed.\n- `family_id` - An identifier for the patent family.\n- `cpc_codes` - A list of the patent classification codes covering the patent.\n\n### **test.csv** → Nearest Neighbor Patents of the 2,500 test patents\n- `publication_number` - Only patents published on or after 1975 were included in this column.\n- `target_[N]` - The Nth nearest neighbor of the patent, based on embeddings from the Google Patents Research Data BigQuery dataset.\n\n### **patent_data/[year_month].parquet** → Patent text from Google Patents Public Data\n- `publication_number` - The patent identifier.\n- `title` - The text of the patent's title.\n- `abstract` - The text of the patent's abstract.\n- `claims` - The text of the patent's claims.\n- `description` - The text of the patent's full description.\n\n### **all_patents.parquet** → Title and Abstract of all the patents compiled from **patent_data/[year_month].parquet** (by @aerdem4)\n- `publication_number` - The patent identifier.\n- `Title` - Title of the patent.\n- `Abstract` - Abstract of the patent.\n\n> **Note:** In this notebook, we will not be using the `description` text of patents to keep the runtime small. Also, we are not be using `patent_data/[year_month].parquet` files as `all_patents.parquet` file contains the necessary information for all the patents. You are welcome to test with all the information of the patents.","metadata":{}},{"cell_type":"code","source":"# Read the CSV file into a DataFrame with specific columns\ntest_df = pd.read_csv(\"/kaggle/input/uspto-explainable-ai/test.csv\")\ntest_df = test_df.iloc[:, :CFG.num_neighbors+1]\ntarget_cols = list(test_df.columns[1:])\n\n# Merge metadata of the patents\nmeta_df = pd.read_parquet(\"/kaggle/input/uspto-explainable-ai/patent_metadata.parquet\")\ntest_df = test_df.merge(meta_df, on=\"publication_number\", how=\"left\")\n\n# Merge Title and Abstract of the patennts\npatent_df = pd.read_parquet(\"/kaggle/input/uspto-all-patents-after-1975/all_patents.parquet\")\ntest_df = test_df.merge(patent_df, on=\"publication_number\")\n\n# Fill NaN values\ntest_df[\"title\"] = test_df[\"title\"].fillna(\"\")\ntest_df[\"abstract\"] = test_df[\"abstract\"].fillna(\"\")\n\n# Merge Title and Abstract of the neighbour patents\nfor i in range(CFG.num_neighbors):\n    test_df = test_df.merge(\n        patent_df,\n        left_on=target_cols[i],\n        right_on=\"publication_number\",\n        how=\"left\",\n        suffixes=(\"\", f\"_{i}\"),\n    )\n\n    # Fill NaN values\n    test_df[f\"title_{i}\"] = test_df[f\"title_{i}\"].fillna(\"\")\n    test_df[f\"abstract_{i}\"] = test_df[f\"abstract_{i}\"].fillna(\"\")\n\n    # Drop extra publication_number column from merges\n    test_df = test_df.drop(columns=[f\"publication_number_{i}\"])\n\n# Reset index order as it will be used later for iteration\ntest_df = test_df.reset_index(drop=True)\n\n# Clean up memory\ndel meta_df, patent_df\ngc.collect()","metadata":{"papermill":{"duration":28.857128,"end_time":"2024-05-07T13:55:16.654504","exception":false,"start_time":"2024-05-07T13:54:47.797376","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-30T08:45:29.200159Z","iopub.execute_input":"2024-05-30T08:45:29.200543Z","iopub.status.idle":"2024-05-30T08:47:16.745365Z","shell.execute_reply.started":"2024-05-30T08:45:29.20051Z","shell.execute_reply":"2024-05-30T08:47:16.744386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2024-06-07T04:58:30.104662Z","iopub.execute_input":"2024-06-07T04:58:30.105102Z","iopub.status.idle":"2024-06-07T04:58:30.467536Z","shell.execute_reply.started":"2024-06-07T04:58:30.105068Z","shell.execute_reply":"2024-06-07T04:58:30.466185Z"},"trusted":true},"execution_count":1,"outputs":[{"traceback":["\u001b[0;31m---------------------------------------------------------------------------\u001b[0m","\u001b[0;31mNameError\u001b[0m                                 Traceback (most recent call last)","Cell \u001b[0;32mIn[1], line 2\u001b[0m\n\u001b[1;32m      1\u001b[0m \u001b[38;5;66;03m# Read the CSV file into a DataFrame with specific columns\u001b[39;00m\n\u001b[0;32m----> 2\u001b[0m test_df \u001b[38;5;241m=\u001b[39m \u001b[43mpd\u001b[49m\u001b[38;5;241m.\u001b[39mread_csv(\u001b[38;5;124m\"\u001b[39m\u001b[38;5;124m/kaggle/input/uspto-explainable-ai/test.csv\u001b[39m\u001b[38;5;124m\"\u001b[39m)\n\u001b[1;32m      3\u001b[0m test_df \u001b[38;5;241m=\u001b[39m test_df\u001b[38;5;241m.\u001b[39miloc[:, :CFG\u001b[38;5;241m.\u001b[39mnum_neighbors\u001b[38;5;241m+\u001b[39m\u001b[38;5;241m1\u001b[39m]\n\u001b[1;32m      4\u001b[0m target_cols \u001b[38;5;241m=\u001b[39m \u001b[38;5;28mlist\u001b[39m(test_df\u001b[38;5;241m.\u001b[39mcolumns[\u001b[38;5;241m1\u001b[39m:])\n","\u001b[0;31mNameError\u001b[0m: name 'pd' is not defined"],"ename":"NameError","evalue":"name 'pd' is not defined","output_type":"error"}]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:47:16.747595Z","iopub.execute_input":"2024-05-30T08:47:16.747875Z","iopub.status.idle":"2024-05-30T08:47:16.774249Z","shell.execute_reply.started":"2024-05-30T08:47:16.747852Z","shell.execute_reply":"2024-05-30T08:47:16.773381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📋 | Solution Overview\n\nBefore we dive deep into the problem, let's have a look at the overview of our solution.\n\n![Solution Overview](https://i.postimg.cc/vHpTVpB0/USPTO-overview.png)\n\nOur approach can be summarized in the following steps:\n\n1. **Generate Prompts**: For each patent, we will generate prompts from all its (patent, neighbor patent) pairs. This involves taking a given patent and pairing it with each of its neighbor patents to form pairs like `[(patent_0, neighbor_0), (patent_0, neighbor_1), ... (patent_0, neighbor_N)]`.\n\n2. **Extract Keywords**: Using a Large Language Model (LLM), we will iteratively extract keywords from each patent pair. This process involves iteratively feeding the LLM with the generated prompts and obtaining relevant keywords from the responses.\n\n3. **Create Queries**: After extracting keywords for each patent, we will create search queries using binary operators. These quries are the targets of this competition which we will submit in `submission.csv`","metadata":{}},{"cell_type":"markdown","source":"# 🔧 | Prompt Engineering\n\nWe will be using the simple prompt template below, followed by the chat template of Gemma, to extract queries/keywords from each patent and its one neighbor. We will iteratively repeat this process (patent, neighbor) pairs to extract more effective queries. You are welcome to explore more advanced prompt templates for better results.\n\n**Prompt Template:**\n```\nTask:\nAnalyze and compare the given two patents and identify the common and similar query keywords that should yield these two patents when searched in the United States Patent and Trademark Office (USPTO) database.\n\nInstructions:\n1. Carefully read and understand the provided 'Patent 1' and 'Patent 2' below. Pay attention to their titles and abstracts.\n2. Identify the key terms, concepts, and components that are common or similar in both patents, that are likely to be used when searching for these patents.\n3. After the 'Keywords' section below, write only the keywords, each separated with a semicolone (';') and a space (' '). For example, 'keyword1; keyword2; keyword3_1 keyword3_2'.\n4. You are not allowed to add any narratives or text before or after your response.\n\nPatent 1:\n* Title: ...\n* Abstract: ...\n\nPatent 2:\n* Title: ...\n* Abstract: ...\n\nKeywords:\n\n```\n\n**Chat Template:**\n\n```\n<start_of_turn>user\nPrompt.....<end_of_turn>\n<start_of_turn>model\n```","metadata":{}},{"cell_type":"code","source":"prompt_template = \"Task:\\nAnalyze and compare the given two patent abstracts and titles, and identify the common or similar query keywords that should yield these two patents when searched in the United States Patent and Trademark Office (USPTO) database.\\n\\nInstructions:\\n1. Carefully read and understand the provided 'Patent 1' and 'Patent 2' titles and abstracts below.\\n2. Identify the key terms, concepts, and components that are either common or similar in both patent titles and abstracts.\\n3. In the 'Keywords' section below, write the common or similar keywords, separating each keyword with a semicolon (';') and a space (' '). Here is an example response, 'keyword1; keyword2; keyword3_1 keyword3_2'.\\n4. Do not add any additional narratives or text before or after the keywords.\\n\\nPatent 1:\\n* Title: {title_a}\\n* Abstract: {abstract_a}\\n\\nPatent 2:\\n* Title: {title_b}\\n* Abstract: {abstract_b}\\n\\nKeywords:\"\n\nchat_template = f\"<start_of_turn>user\\n{prompt_template}<end_of_turn>\\n<start_of_turn>model\\n\"","metadata":{"execution":{"iopub.status.busy":"2024-05-30T18:45:55.762487Z","iopub.execute_input":"2024-05-30T18:45:55.763179Z","iopub.status.idle":"2024-05-30T18:45:55.767777Z","shell.execute_reply.started":"2024-05-30T18:45:55.763145Z","shell.execute_reply":"2024-05-30T18:45:55.766853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(prompt_template)","metadata":{"execution":{"iopub.status.busy":"2024-05-30T18:30:44.804463Z","iopub.execute_input":"2024-05-30T18:30:44.804713Z","iopub.status.idle":"2024-05-30T18:30:44.809295Z","shell.execute_reply.started":"2024-05-30T18:30:44.804692Z","shell.execute_reply":"2024-05-30T18:30:44.808475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_prompt(row, neighbor_idx):\n    prompt = chat_template.format(title_a=row[\"title\"], abstract_a=row[\"abstract\"],\n                                  title_b=row[f\"title_{neighbor_idx}\"], abstract_b=row[f\"abstract_{neighbor_idx}\"])\n    return prompt","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:48:04.606537Z","iopub.execute_input":"2024-05-30T08:48:04.60725Z","iopub.status.idle":"2024-05-30T08:48:04.611903Z","shell.execute_reply.started":"2024-05-30T08:48:04.607217Z","shell.execute_reply":"2024-05-30T08:48:04.610945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check Sample Prompt","metadata":{}},{"cell_type":"code","source":"prompt_sample = create_prompt(test_df.iloc[2], 0)\nprint(prompt_sample)","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:48:19.26196Z","iopub.execute_input":"2024-05-30T08:48:19.262576Z","iopub.status.idle":"2024-05-30T08:48:19.267853Z","shell.execute_reply.started":"2024-05-30T08:48:19.262543Z","shell.execute_reply":"2024-05-30T08:48:19.266909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤖 | Modeling\n\nTo extract the query keywords from patents either we can use simple text processing/filtering but it is very difficult to generalize the filter for such keywords. On the other side, we can use methods like Named Entitry Recogntition (NER) to identfy the key terms in the patent but int this competition we don't have any ground truth query keywords in the first place thus it not possible to train a NER model for this task. Finally, we can use large language model (LLM) which can follow instruction based on our prompt for effectively extracting query keywords from patents.\n\nIn this notebook, we will be using **Gemma 1.1 2b instructed tuned** model from **KerasNLP**. KerasNLP has many more recent and powerful pretrained LLMs. To name a few, `Phi3`, `Llama3`, `Mistral`, etc. You are welcome to try them out. You can check the available pretrained LLMs in the [KerasNLP webpage](https://keras.io/api/keras_nlp/models/).`\n\n> We are using the \"Instruction tuned\" model instead of the \"Pretrained\" one because it is easier for the model to fine-tune on the prepared dataset. In the ","metadata":{}},{"cell_type":"code","source":"# Declare the model\ngemma_lm = keras_nlp.models.GemmaCausalLM.from_preset(\"gemma_1.1_instruct_2b_en\")\n\n# Set input length of small to keep the memory and latency cost small\ngemma_lm.preprocessor.sequence_length = CFG.input_length","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:48:32.142138Z","iopub.execute_input":"2024-05-30T08:48:32.142542Z","iopub.status.idle":"2024-05-30T08:49:34.017138Z","shell.execute_reply.started":"2024-05-30T08:48:32.142511Z","shell.execute_reply":"2024-05-30T08:49:34.01607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🔑 | Keyword Generation\n\nThe following code will extract query keywords from each patent and one of its neighbors. First, it will use Gemma LLM to analyze the patents and try to identify the common query keywords. Then, we will apply text processing to extract the query keywords from the model's output. We will repeat this process $N$ times as defined in `num_neighbors` of the config file. Finally, we will merge the keywords for all the (patent, neighbor) pair. In this notebook, we will use $2$ neighbors. You are welcome to use more neighbor patents.\n\n> **Note:** In the code for extracting keywords from the model output, we have included several text processing steps which may seem unnecessary. However, these processes are added to handle various possible pitfalls that could result in `submission scoring errors`. Because, sometimes, LLMs do not follow instructions properly, creating unformatted outputs. Thus, the following code will ensure that we get our query keywords properly from the model output even when the model fails to follow instructions properly.\n","metadata":{}},{"cell_type":"code","source":"def generate_keyword(row, neighbor_idx):\n    # Check if any title or abstract is an empty string\n    fields = [\n        \"title\",\n        \"abstract\",\n        f\"title_{neighbor_idx}\",\n        f\"abstract_{neighbor_idx}\",\n    ]\n    if all(row[field] == \"\" for field in fields):\n        return [\"\"]\n\n    # Create prompt\n    prompt = create_prompt(row, neighbor_idx)\n\n    try:\n        # Generate output from model\n        output = gemma_lm.generate(prompt, max_length=CFG.output_length)\n\n        # Extract keyword from model output\n        keyword = decode_output(output, prompt)\n    except:\n        keyword = [\"\"]\n        \n    return keyword\n\n\ndef decode_output(output, prompt):\n    # Remove input prompt from model output\n    answer = output.replace(prompt, \"\").strip()\n\n    # Avoid edge case when output_max_length < model_output\n    if \"Title:\" in answer and \"Abstract:\" in answer:\n        return [\"\"]\n\n    # Filter out possible unwanted output text\n    for x in [\"Keywords:\", \"solution:\", \"Solution\", \"**\", \"\\n\\n\", \"[\", \"]\"]:\n        answer = answer.replace(x, \"\").strip()\n\n    # Create list of keywords using possible delimiters\n    for sep in [\";\\n-\", \",\\n-\", \"\\n-\", \";\\n\", \";\\n*\", \",\\n*\", \",\\n\", \"\\n*\", \",\", \";\"]:\n        if sep in answer:\n            answer = answer.strip(sep).strip().split(sep + \" \")\n\n    # Final filtering: remove '*', '.', and any keywords length > 40\n    keywords = [x.replace(\"*\", \"\").replace(\".\", \"\") for x in set(answer) if len(x) < 40]\n    \n    # If there is no keywords found, then enter a empty string as keyword\n    if not len(keywords):\n        keywords = [\"\"]\n        \n    return keywords\n","metadata":{"papermill":{"duration":0.017948,"end_time":"2024-05-07T13:59:31.901066","exception":false,"start_time":"2024-05-07T13:59:31.883118","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-30T08:50:09.31842Z","iopub.execute_input":"2024-05-30T08:50:09.318797Z","iopub.status.idle":"2024-05-30T08:50:09.329705Z","shell.execute_reply.started":"2024-05-30T08:50:09.318766Z","shell.execute_reply":"2024-05-30T08:50:09.328668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's generate the query keywords from the patents.","metadata":{}},{"cell_type":"code","source":"# Keywords for all patents\nkeywords_all = []\n\nfor i, row in tqdm(test_df.iterrows(), total=test_df.shape[0]):\n    # Keywords for one patent\n    keywords = []\n    \n    # Iteratively create keywords for each (patent, neighbour) pair\n    for i in range(CFG.num_neighbors):\n        keywords += generate_keyword(row, neighbor_idx=i)\n        \n    # Remove duplicate keywords\n    keywords = list(set(keywords))\n    \n    # Merge keywords\n    keywords_all.append(keywords)","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:50:23.239927Z","iopub.execute_input":"2024-05-30T08:50:23.240845Z","iopub.status.idle":"2024-05-30T08:50:53.959965Z","shell.execute_reply.started":"2024-05-30T08:50:23.24081Z","shell.execute_reply":"2024-05-30T08:50:53.958958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check some of the keywords that we extracted from patents.","metadata":{}},{"cell_type":"code","source":"_ = [print(f\"Keywords {i}: {q}\", end=\"\\n\\n\") for i, q in  enumerate(keywords_all[:3])]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-30T08:55:59.63704Z","iopub.execute_input":"2024-05-30T08:55:59.637976Z","iopub.status.idle":"2024-05-30T08:55:59.643159Z","shell.execute_reply.started":"2024-05-30T08:55:59.637934Z","shell.execute_reply":"2024-05-30T08:55:59.642262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🚫 | Whoosh Filter\n\nNow that we have the keywords, we will soon create search queries from them. During submission, these queries will be used in the patent search database emulator, built with [Whoosh](https://whoosh.readthedocs.io/en/latest/intro.html) for searching patents. The returned patents from the search will be used for scoring. In simple terms, we can consider Whoosh as a search engine like Google, where given an input query, it returns the patents that match those queries. However, unlike Google, we can't input just anything into Whoosh; we need to maintain a specific format. To ensure our input works properly in Whoosh, we need to filter the query. The following code will remove common words (stopwords) and numbers (e.g., `123`, `1,234`, or `1.234`) from the keywords and make them searchable in Whoosh.\n\nYou can learn more about the Whoosh filter [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499582). If you want to learn more about how search in Whoosh works, check out this [demo](https://www.kaggle.com/code/sohier/basic-whoosh-search-demo).\n","metadata":{}},{"cell_type":"code","source":"BRS_STOPWORDS = ['an', 'are', 'by', 'for', 'if', 'into', 'is', 'no', 'not', 'of', 'on', 'such',\n        'that', 'the', 'their', 'then', 'there', 'these', 'they', 'this', 'to', 'was', 'will', 'and', 'or']\nNUMBER_REGEX = re.compile(r'^(\\d+|\\d{1,3}(,\\d{3})*)(\\.\\d+)?$')\n\nclass NumberFilter(whoosh.analysis.Filter):\n    def __call__(self, tokens):\n        for t in tokens:\n            if not NUMBER_REGEX.match(t.text):\n                yield t\n\ncustom_analyzer = whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-30T08:51:37.020967Z","iopub.execute_input":"2024-05-30T08:51:37.021355Z","iopub.status.idle":"2024-05-30T08:51:37.028308Z","shell.execute_reply.started":"2024-05-30T08:51:37.021323Z","shell.execute_reply":"2024-05-30T08:51:37.027499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see this filter in action. It is apparent that this filter removes numbers and stopwords.\n","metadata":{}},{"cell_type":"code","source":"it = custom_analyzer(\"device, 1.023, machine, that, learning, there\")\n[token.text for token in it]","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:53:30.632367Z","iopub.execute_input":"2024-05-30T08:53:30.633086Z","iopub.status.idle":"2024-05-30T08:53:30.639341Z","shell.execute_reply.started":"2024-05-30T08:53:30.633051Z","shell.execute_reply":"2024-05-30T08:53:30.638357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the final step of the filtering, we will do a check using `QueryValidator()` from `whoosh_utils` library to ensure the query is fit to be used in search.\n\n> **Note:** Whoosh library itself is a bit complicated thus this competition offers `whoosh_utils` which allows us to search, validate queries more easily.","metadata":{}},{"cell_type":"code","source":"query_validator = whoosh_utils.QueryValidator()\n\ndef validate_query(query):\n    query = \"ti:device\" if not len(query) or not isinstance(query, str) else query\n    try:\n        query_validator.validate_query(query)\n    except:\n        query = \"ti:device\"\n    return query","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-30T08:53:39.76862Z","iopub.execute_input":"2024-05-30T08:53:39.76902Z","iopub.status.idle":"2024-05-30T08:53:39.785129Z","shell.execute_reply.started":"2024-05-30T08:53:39.768989Z","shell.execute_reply":"2024-05-30T08:53:39.784208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"validate_query(\"device OR machine\") # query is valid","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:53:41.691279Z","iopub.execute_input":"2024-05-30T08:53:41.692173Z","iopub.status.idle":"2024-05-30T08:53:41.698397Z","shell.execute_reply.started":"2024-05-30T08:53:41.692128Z","shell.execute_reply":"2024-05-30T08:53:41.697524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"validate_query(\"(device OR machine\") # query is invalid due to missing ')' thus returns default query","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:53:45.303011Z","iopub.execute_input":"2024-05-30T08:53:45.304003Z","iopub.status.idle":"2024-05-30T08:53:45.309762Z","shell.execute_reply.started":"2024-05-30T08:53:45.303968Z","shell.execute_reply":"2024-05-30T08:53:45.308767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🔍 | Query Creation\n\n## What are the Queries?\n\nThe search queries here are the target of the competition. Specifically, they are binary queries, composed using binary logic operators to evaluate the truthfulness of each part. These queries allow for complex search criteria, refining search results effectively.\n\nWe combine keywords using different Boolean or Proximity operators:\n- Boolean: `OR`, `AND`, `NOT`, `XOR`\n- Proximity: `ADJ`, `NEAR` (Note: Proximity operators aren't supported for the CPC field.)\n\nWildcards can also be used to include variations of the keywords:\n- Wildcards: `*`, `?`, `$` (Note: Wildcards are incompatible with proximity operators and may impact performance.)\n\nThe following fields of the patents are searchable using the query:\n- `ti`: title\n- `ab`: abstract\n- `clm`: claims\n- `detd`: description\n- `cpc`: CPC codes\n\n## How are queries created?\n\nIn the following code, we first create a query from CPC codes (up to 15) using binary operators. Then, we create a query from keywords extracted from patents using LLM. We merge queries from both keywords and CPC codes using boolean operators, specifying which query belongs to which field (`ti`, `detd`, or `cpc`). Finally, we perform a query token count check as the competition metric allows only 50 tokens. We also validate the query to ensure everything is in order.\n\nExample queries may look like:\n- `\"detd:(device OR machine)\"`\n- `\"cpc:(C12Q1/485 OR G01N2570/00)\"`\n- `\"detd:(device OR machine) AND cpc:(C12Q1/485 OR G01N2570/00)\"`\n- `\"ti:device\"`\n\n> **Note**: Although query keywords are extracted from titles and abstracts of patents, we search in the description of patents as titles and abstracts might not always contain the keyword. Moreover, keywords appearing in titles and abstracts are very likely to appear in the description of the patent.","metadata":{}},{"cell_type":"code","source":"queries = []\n\nfor i, row in tqdm(test_df.iterrows(), total=len(test_df)):\n    # Create query from cpc_codes\n    cpc = row[\"cpc_codes\"]\n    query_cpc = f\"cpc:({' OR '.join(cpc[:15])})\" if len(cpc) else \"\"\n    \n    try:\n        # Analyze the keywords\n        keywords_str = \", \".join(keywords_all[i])\n        tokens = list(set([token.text for token in custom_analyzer(keywords_str)]))\n        \n        # Reduce the keywords if number of query tokens > 50\n        while len(tokens):\n            # Create query from keywords\n            query_keywords = f\"({' OR '.join(tokens)})\"\n            \n            # Merge quries from keywords and cpc_codes\n            query_check = f\"detd:{query_keywords}\" + (f\" AND {query_cpc}\" if len(query_cpc) else \"\")\n            \n            # Return query if number of query tokens is okay\n            if whoosh_utils.count_query_tokens(query_check) < 50:\n                query = query_check\n                break\n                \n            # Reduce keywords if number query is not okay\n            tokens.pop()\n    except:\n        query = query_cpc\n    \n    # Final query validation\n    query = validate_query(query)\n    \n    queries.append(query)","metadata":{"execution":{"iopub.status.busy":"2024-05-30T08:53:55.947842Z","iopub.execute_input":"2024-05-30T08:53:55.948248Z","iopub.status.idle":"2024-05-30T08:53:55.963746Z","shell.execute_reply.started":"2024-05-30T08:53:55.948219Z","shell.execute_reply":"2024-05-30T08:53:55.962704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check some of the quries we created.","metadata":{}},{"cell_type":"code","source":"_ = [print(f\"Query {i}: {q}\", end=\"\\n\\n\") for i, q in  enumerate(queries[:3])]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-30T08:55:30.230455Z","iopub.execute_input":"2024-05-30T08:55:30.230853Z","iopub.status.idle":"2024-05-30T08:55:30.236786Z","shell.execute_reply.started":"2024-05-30T08:55:30.230824Z","shell.execute_reply":"2024-05-30T08:55:30.235807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📬 | Submission\n\nFollowing code will prepare the submission file.","metadata":{}},{"cell_type":"code","source":"test_df[\"query\"] = queries\npred_df = test_df[[\"publication_number\", \"query\"]]\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"papermill":{"duration":0.022304,"end_time":"2024-05-07T13:59:50.358151","exception":false,"start_time":"2024-05-07T13:59:50.335847","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-30T08:56:25.265257Z","iopub.execute_input":"2024-05-30T08:56:25.26617Z","iopub.status.idle":"2024-05-30T08:56:25.283817Z","shell.execute_reply.started":"2024-05-30T08:56:25.266124Z","shell.execute_reply":"2024-05-30T08:56:25.28293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"But as this competition is quite different from other NLP competitions due to its weakly supervised labels, let's try to understand how the submission process works to have a better understanding of the competition.\n\nOnce submitted, the `whoosh_utils` library will search patents using queries from our `submission.csv`. Once it has retrieved the patents, they will be evaluated using mean average precision at 50 (`mAP@50`) between the retrieved patents and the ground truth patents. So, unlike other Kaggle competitions where the submission file is directly used for scoring, here the Whoosh library will be used before scoring.\n","metadata":{}},{"cell_type":"markdown","source":"# 🔭 | Future Directions\n\nLooking forward, we can further improve this notebook by:\n\n1. Trying different models like `Phi3`, `Llama3`, or `Mistral`.\n2. Increasing `num_neighbors` in the config file to extract more keywords.\n3. Using fewer `cpc` codes in the query (currently up to 15) to include more query keywords.\n4. Experimenting with different operator combinations. Currently, we are using simple `OR` and `AND`.\n5. Utilizing more advanced prompt engineering.\n","metadata":{}},{"cell_type":"markdown","source":"# 📌 | Reference\n\n* [USPTO LLM Solution [Prompt Eng]](https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng)\n* [USPTO: whoosh==2.7.5 Offline Use](https://www.kaggle.com/code/seshurajup/uspto-whoosh-2-7-5-offline-use)","metadata":{}}]}