{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59575,"databundleVersionId":8060720,"sourceType":"competition"},{"sourceId":7731449,"sourceType":"datasetVersion","datasetId":4517815},{"sourceType":"kernelVersion","sourceId":174185912}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# About the competition - \n>From The Description - The competition aims to create a model that generates Boolean search queries for patent documents. This model should produce queries that retrieve the same patents as the input set. Successful solutions will enhance AI-powered patent searches, making results more interpretable for professionals and facilitating responsible AI adoption in intellectual property.\n# My Understanding- \n>In this competition we have to create a model that can predict a query that can fectch the exact data we need based on the details about the query.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"markdown","source":"# Patent Data\n>patent_data/[year_month].parquet Patent text from Google Patents Public Data by IFI CLAIMS Patent Services and Google. The competition data is current through July, 2023.\n\n* **publication_number**\n* **title** - The text of the patent's title.\n* **abstract** - The text of the patent's abstract.\n* **claims** - The text of the patent's claims.\n* **description** - The text of the patent's full description.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\npatent_data_sample = pd.read_parquet(\"/kaggle/input/uspto-explainable-ai/patent_data/1836_11.parquet\")\ndisplay(patent_data_sample.head(10))","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:05:36.435258Z","iopub.execute_input":"2024-05-01T07:05:36.435590Z","iopub.status.idle":"2024-05-01T07:05:36.995673Z","shell.execute_reply.started":"2024-05-01T07:05:36.435543Z","shell.execute_reply":"2024-05-01T07:05:36.994286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Patent Metadata\n>Patent_metadata.parquet Metadata for the most recent patents in each patent family.\n\n* **publication_number** - The patent identifier.\n* **publication_date** - The date the patent was published.\n* **filing_date** - The date the patent was filed.\n* **family_id** - An identifier for the patent family.\n* **cpc_codes** - A list of the cooperative patent classifcation codes covering the patent.\n","metadata":{}},{"cell_type":"code","source":"patent_metadata = pd.read_parquet(\"/kaggle/input/uspto-explainable-ai/patent_metadata.parquet\")\ndisplay(patent_metadata.head(10))","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:05:36.997531Z","iopub.execute_input":"2024-05-01T07:05:36.998096Z","iopub.status.idle":"2024-05-01T07:06:03.059730Z","shell.execute_reply.started":"2024-05-01T07:05:36.998063Z","shell.execute_reply":"2024-05-01T07:06:03.058626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Nearest Neighbour\n>nearest_neighbors.csv The 50 patents most similar to the target patent.\n\n* **publication_number**\n* **neighbor_[N]** - The Nth nearest neighbor of the patent, based on embeddings from the Google Patents Research Data BigQuery dataset by IFI CLAIMS Patent Services and Google. The competition data is current through July, 2023.\n> This file is very large so I have read only first 100 rows.","metadata":{}},{"cell_type":"code","source":"neighbour = pd.read_csv(\"/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv\", nrows = 100)\ndisplay(neighbour.head(10))","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:06:03.062377Z","iopub.execute_input":"2024-05-01T07:06:03.062798Z","iopub.status.idle":"2024-05-01T07:06:03.105951Z","shell.execute_reply.started":"2024-05-01T07:06:03.062762Z","shell.execute_reply":"2024-05-01T07:06:03.104914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Range of Data","metadata":{}},{"cell_type":"code","source":"import os\npatent_data_path = \"/kaggle/input/uspto-explainable-ai/patent_data\"\nyears = []\nfor i in os.listdir(patent_data_path):\n    try:\n        year = int(i.split(\"_\")[0])\n        years.append(year)\n    except:\n        print()\nyears = list(set(years))       \noldest = min(years)\nlatest = max(years)\nduration = len(years)\nprint(\"We have data from year \"+str(oldest)+\" to \"+str(latest)+\". So we have patent data of \"+str(duration)+\" years.\")\n","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:06:03.107202Z","iopub.execute_input":"2024-05-01T07:06:03.107611Z","iopub.status.idle":"2024-05-01T07:06:03.382315Z","shell.execute_reply.started":"2024-05-01T07:06:03.107545Z","shell.execute_reply":"2024-05-01T07:06:03.381393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Index\n>A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975. The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook. The following fields are searchable:\n\n* **ti**: title.\n* **ab**: abstract.\n* **clm**: claims.\n* **detd**: description.\n* **cpc**: cpc codes.\n\n**What is Whoosh?**\n>Whoosh is a fast, pure Python search engine library.\nThe primary design impetus of Whoosh is that it is pure Python. You should be able to use Whoosh anywhere you can use Python, no compiler or Java required.\nLike one of its ancestors, Lucene, Whoosh is not really a search engine, it’s a programmer library for creating a search engine.\nPractically no important behavior of Whoosh is hard-coded. Indexing of text, the level of information stored for each term in each field, parsing of search queries, the types of queries allowed, scoring algorithms, etc. are all customizable, replaceable, and extensible.\n\nThe authors created a notebook on how to use whoosh. you can check their notebook [here](https://www.kaggle.com/code/sohier/basic-whoosh-search-demo)","metadata":{}},{"cell_type":"markdown","source":"# Using Whoosh\n","metadata":{}},{"cell_type":"code","source":"import whoosh_utils","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:06:03.383589Z","iopub.execute_input":"2024-05-01T07:06:03.383870Z","iopub.status.idle":"2024-05-01T07:06:03.388279Z","shell.execute_reply.started":"2024-05-01T07:06:03.383848Z","shell.execute_reply":"2024-05-01T07:06:03.387075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_idx = whoosh_utils.load_index('/kaggle/input/uspto-explainable-ai/train_index')\nsearcher = whoosh_utils.get_searcher(train_idx)\nqp = whoosh_utils.get_query_parser()","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:06:03.389851Z","iopub.execute_input":"2024-05-01T07:06:03.390764Z","iopub.status.idle":"2024-05-01T07:06:09.299522Z","shell.execute_reply.started":"2024-05-01T07:06:03.390727Z","shell.execute_reply":"2024-05-01T07:06:09.298346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"query = 'ti:balloons OR ti:string'\nresult = whoosh_utils.execute_query(query, qp, searcher)[:5]\nresult","metadata":{"execution":{"iopub.status.busy":"2024-05-01T07:06:09.300705Z","iopub.execute_input":"2024-05-01T07:06:09.303046Z","iopub.status.idle":"2024-05-01T07:06:09.446234Z","shell.execute_reply.started":"2024-05-01T07:06:09.303011Z","shell.execute_reply":"2024-05-01T07:06:09.445225Z"},"trusted":true},"execution_count":null,"outputs":[]}]}