{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59575,"databundleVersionId":8060720,"sourceType":"competition"},{"sourceId":8220803,"sourceType":"datasetVersion","datasetId":4873590},{"sourceId":167206324,"sourceType":"kernelVersion"}],"dockerImageVersionId":30702,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on April 24, 2024. By Marília Prata, mpwolke","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport plotly.graph_objs as go\nimport plotly.offline as py\nimport plotly.express as px\n\n#Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-25T00:06:59.238813Z","iopub.execute_input":"2024-04-25T00:06:59.239426Z","iopub.status.idle":"2024-04-25T00:07:01.411154Z","shell.execute_reply.started":"2024-04-25T00:06:59.239384Z","shell.execute_reply":"2024-04-25T00:07:01.408793Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Competition Citation:\n\n@misc{uspto-explainable-ai,\n\n    author = {Scott Beliveau, Don Cenkci, Diego Alcantara, Keith Frost, Don Garret, Kiran Reddy, Sohier Dane, Will Cukierski, Addison Howard, Maggie Demkin},\n    \n    title = {USPTO - Explainable AI for Patent Professionals},\n    \n    publisher = {Kaggle},\n    year = {2024},\n    url = {https://kaggle.com/competitions/uspto-explainable-ai}\n}\n\nUSPTO (United States Patent and Trademark Office)\n\n\"The United States Patent and Trademark Office (USPTO) is the federal agency for granting U.S. patents and registering trademarks.\" \n\nThe USPTO offers one of the largest repositories of scientific, technical, and commercial information in the world through its Open Data Portal. Patents are a form of intellectual property granted in exchange for the public disclosure of new and useful inventions. Because patents undergo an intensive vetting process prior to grant, and because the history of U.S. innovation spans over two centuries and 11 million patents, the U.S. patent archives stand as a rare combination of data volume, quality, and diversity.\"\n\nhttps://www.kaggle.com/competitions/uspto-explainable-ai","metadata":{}},{"cell_type":"markdown","source":"![](https://s7280.pcdn.co/wp-content/uploads/2020/07/interpritability.png)https://www.bmc.com/blogs/machine-learning-interpretability-vs-explainability/","metadata":{}},{"cell_type":"markdown","source":"#Interpretability of Machine Learning\n\nInterpretability of Machine Learning: Recent Advances and Future Prospects\n\nAuthors: Gao, Lei and Guan, Ling\n\n\"The proliferation of machine learning (ML) has drawn unprecedented interest in the study of  various multimedia contents such as text, image, audio and video, among others. Consequently,understanding and learning ML-based representations have taken center stage in knowledge discovery in intelligent multimedia research and applications. Nevertheless, the black-box nature of contemporary ML, especially in deep neural networks (DNNs), has posed a primary challenge for ML-based representation learning. To address this black-box problem, the studies on interpretability of ML have attracted tremendous interests in recent years.\"\n\n\"This paper presents a survey on recent advances and future prospects on  interpretability of ML, with several application examples pertinent to multimedia computing, including text-image cross-modal representation learning, face recognition, and the recognition of objects. It is evidently shown that the study of interpretability of ML promises an important research direction, one which is worth further investment in.\"\n\nAlgorithm Unrolling\n\n\"Algorithm unrolling solves model interpretability by providing a concrete and systematic connection between iterative algorithms that are widely used in signal process ing and DNNs. Given an iterative algorithm, a corresponding deep network is generated  by cascading its iterations h. Then, iteration step  h is executed a number of times, resulting in different parameters h1, h2,... Each iteration h depends on algorithm parameters, which are transferred into network parameters 1, 2,... Instead of determining parameters through cross-validation or analytical derivations, the parameters 1, 2,... are learned from training datasets through end-to-end training. In this way, the network layers naturally inherit interpretability from the iteration procedure.\"\n\nhttps://arxiv.org/pdf/2305.00537.pdf","metadata":{}},{"cell_type":"markdown","source":"#train_index_patent_ids.json A list of the patents included in the Whoosh index.\n\nhttps://www.kaggle.com/competitions/uspto-explainable-ai/data","metadata":{}},{"cell_type":"code","source":"df1= pd.read_json('../input/uspto-explainable-ai/train_index_patent_ids.json', lines=True)\ndf1.head()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-24T20:38:21.214869Z","iopub.execute_input":"2024-04-24T20:38:21.218033Z","iopub.status.idle":"2024-04-24T20:39:23.568569Z","shell.execute_reply.started":"2024-04-24T20:38:21.217970Z","shell.execute_reply":"2024-04-24T20:39:23.567251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#nearest_neighbors.csv The 50 patents most similar to the target patent.","metadata":{}},{"cell_type":"code","source":"df2= pd.read_csv('../input/uspto-explainable-ai/nearest_neighbors.csv', nrows=1000)\ndf2.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-24T20:46:28.881762Z","iopub.execute_input":"2024-04-24T20:46:28.882240Z","iopub.status.idle":"2024-04-24T20:46:29.198160Z","shell.execute_reply.started":"2024-04-24T20:46:28.882207Z","shell.execute_reply":"2024-04-24T20:46:29.196789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Open One parquet file: Year 2013, Month December  2013_12.parquet","metadata":{}},{"cell_type":"code","source":"#By Bachir https://www.kaggle.com/code/bachrr/gemma2b-finetuned-lora-text2sql\n\ndf = pd.read_parquet(\"/kaggle/input/uspto-explainable-ai/patent_data/2013_12.parquet\")\ndf.tail()","metadata":{"execution":{"iopub.status.busy":"2024-04-24T20:58:16.141430Z","iopub.execute_input":"2024-04-24T20:58:16.141911Z","iopub.status.idle":"2024-04-24T20:58:28.699984Z","shell.execute_reply.started":"2024-04-24T20:58:16.141880Z","shell.execute_reply":"2024-04-24T20:58:28.698665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Interpretability: Definitions, Methods and Applications\n\nInterpretable machine learning: definitions, methods, and applications\n\nAuthors:  W. James Murdocha, Chandan Singhb, Karl Kumbiera, Reza Abbasi-Aslb, and Bin Yua\n\n\"In the absence of a well-formed definition of interpretability, a broad range of methods with a correspondingly broad range of outputs (e.g. visualizations, natural language, mathematical\n equations) have been labeled as interpretation. This has led to considerable confusion about the notion of interpretability. In particular, it is unclear what it means to interpret some thing, what common threads exist among disparate methods, and how to select an interpretation method for a particular problem/audience.\"\n \n Defining interpretable machine learning.\n \n \"On its own, interpretability is a broad, poorly defined concept. Taken to its  full generality, to interpret data means to extract information  (of some form) from it. The set of methods falling under this  umbrella spans everything from designing an initial experiment to visualizing final results. In this overly general form, interpretability is not substantially different from the established concepts of data science and applied statistics.\"\n \nPost hoc analysis\n\n\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question. The process of analyzing the model often involves using interpretability methods to extract various (stable) forms of information from the model. The extracted information can then be analyzed and displayed using standard data analysis methods, such as scatter plots and histograms. The ability of the interpretations to properly describe what the model has learned is denoted by descriptive accuracy.\"\n\nDemonstrating relevancy to real-world problems. \n\n\"Another angle for developing improved interpretation methods is to improve the relevancy of interpretations for some audience or problem. This is normally done by introducing a novel form\n of output, such as feature heatmaps, rationales, feature hierarchies or identifying important elements in the training set.\" \n \nhttps://arxiv.org/pdf/1901.04592.pdf ","metadata":{"execution":{"iopub.status.busy":"2024-04-24T23:37:46.612477Z","iopub.execute_input":"2024-04-24T23:37:46.613005Z","iopub.status.idle":"2024-04-24T23:37:46.623804Z","shell.execute_reply.started":"2024-04-24T23:37:46.612969Z","shell.execute_reply":"2024-04-24T23:37:46.622223Z"}}},{"cell_type":"markdown","source":"#Trying to understand claims.","metadata":{}},{"cell_type":"code","source":"df['claims'][2]","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:49:14.860684Z","iopub.execute_input":"2024-04-24T21:49:14.861215Z","iopub.status.idle":"2024-04-24T21:49:14.879078Z","shell.execute_reply.started":"2024-04-24T21:49:14.861178Z","shell.execute_reply":"2024-04-24T21:49:14.877333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.utils import io\nwith io.capture_output() as captured:\n    !pip install scispacy\n    !pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.2.4/en_core_sci_lg-0.2.4.tar.gz\n    !pip install whoosh","metadata":{"execution":{"iopub.status.busy":"2024-04-24T20:58:47.728172Z","iopub.execute_input":"2024-04-24T20:58:47.728648Z","iopub.status.idle":"2024-04-24T21:04:15.997517Z","shell.execute_reply.started":"2024-04-24T20:58:47.728615Z","shell.execute_reply":"2024-04-24T21:04:15.994657Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing\n\nAuthors: Vishal Monga, Senior Member, Yuelong Li and Yonina C. Eldar\n\n\"Deep neural networks provide unprecedented performance gains in many real world problems in signal and\n image processing. Despite these gains, future development and practical deployment of deep networks is hindered by their black box nature (i.e.lack of interpretability, and by the need for very large training sets).\"\n\n\"An emerging technique called algorithm unrolling or unfolding offers promise in eliminating these issues (i.e.lack of interpretability, and by the need for very large training sets) by providing a concrete and systematic connection between iterative algorithms that are used widely in signal processing and deep neural networks.\"\n\n\"Unrolling methods were first proposed to develop fast neural network approximations for sparse coding. More recently, this direction has attracted enormous attention and is rapidly growing both in theoretic investigations and practical applications. The growing popularity of unrolled deep  networks is due in part to their potential in developing efficient, high-performance and yet interpretable network architectures from reasonable size training sets.\"\n\n\"In this article, the authors reviewed algorithm unrolling for signal and image processing. They extensively covered popular techniques for algorithm unrolling in various domains of signal and image processing including imaging, vision and recognition, and speech processing. By reviewing previous works, they revealed the connections between iterative algorithms and neural networks and present recent theoretical results. Finally, the authors provided a discussion on current limitations of unrolling and suggest possible future research directions.\" \n\nUnrolling Sparse Coding Algorithms into Deep Networks\n\n\"The earliest work in algorithm unrolling dates back to Gregor et al.’s paper (2010) on improving the computational efficiency of sparse coding algorithms through end-to-end training. In particular, they discussed how to improve the efficiency of the Iterative Shrinkage and Thresholding Algorithm (ISTA), one of the most popular approaches in sparse coding. Learned ISTA. Each iteration of ISTA comprises one linear operation followed by a non-linear soft-thresholding operation, which mimics the ReLU activation function. A diagram representation of one iteration step reveals its resemblance to network is dubbed Learned ISTA (LISTA).\"\n\nBridging the Gap between Theory and Practice: \n\n\"While substantial progress has been achieved towards understanding the network behavior through unrolling, more works need to be done to thoroughly understand its mechanism. Although the\n effectiveness of some networks on image reconstruction tasks has been explained somehow by drawing parallels to sparse coding algorithms, it is still mysterious why state-of-the art networks perform well on various recognition tasks. Further more, unfolding itself is not uniquely defined. For instance, there are multiple ways to choose the underlying iterative algorithms, to decide what parameters become trainable and what parameters to fix, and more.\"\n \n\"Another interesting direction is to develop a theory that provides guidance for practical applications. For instance, it is interesting to perform analysis that guide practical network\ndesign choices, such as dimensions of parameters, network depth, etc. It is particularly interesting to identify factors that have high impact on network performance.\"\n\nhttps://arxiv.org/pdf/1912.10557.pdf","metadata":{}},{"cell_type":"code","source":"from collections import defaultdict\n\nimport whoosh\nfrom whoosh.qparser import *\nfrom whoosh.fields import Schema, TEXT, KEYWORD, ID, STORED\nfrom whoosh.analysis import StandardAnalyzer\nfrom whoosh import index\n\nfrom whoosh.analysis import Tokenizer, Token\nfrom whoosh import highlight","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:05:06.472437Z","iopub.execute_input":"2024-04-24T21:05:06.473295Z","iopub.status.idle":"2024-04-24T21:05:06.580634Z","shell.execute_reply.started":"2024-04-24T21:05:06.473216Z","shell.execute_reply":"2024-04-24T21:05:06.579422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.core.display import display, HTML","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:05:11.206286Z","iopub.execute_input":"2024-04-24T21:05:11.206878Z","iopub.status.idle":"2024-04-24T21:05:11.215702Z","shell.execute_reply.started":"2024-04-24T21:05:11.206835Z","shell.execute_reply":"2024-04-24T21:05:11.213763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install scispacy","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:13:39.390241Z","iopub.execute_input":"2024-04-24T21:13:39.391929Z","iopub.status.idle":"2024-04-24T21:13:57.150937Z","shell.execute_reply.started":"2024-04-24T21:13:39.391870Z","shell.execute_reply":"2024-04-24T21:13:57.149026Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Install  en_core_sci_lg","metadata":{}},{"cell_type":"code","source":"#https://github.com/allenai/scispacy/blob/main/scripts/install_remote_packages.py\n\nimport os\n\nfrom scispacy.version import VERSION\n\n\ndef main():\n    s3_prefix = \"https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/\"\n    model_names = [\n        \"en_core_sci_sm\",\n        \"en_core_sci_lg\"\n        \n    ]\n\n    full_package_paths = [\n        f\"{s3_prefix}{model_name}-{VERSION}.tar.gz\" for model_name in model_names\n    ]\n\n    for package_path in full_package_paths:\n        os.system(f\"pip install {package_path}\")\n\n\nif __name__ == \"__main__\":\n    main()","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:29:20.124900Z","iopub.execute_input":"2024-04-24T21:29:20.125355Z","iopub.status.idle":"2024-04-24T21:31:41.603693Z","shell.execute_reply.started":"2024-04-24T21:29:20.125320Z","shell.execute_reply":"2024-04-24T21:31:41.601888Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Stopwords\n\nhttps://ppubs.uspto.gov/pubwebapp/static/pages/stopwords.html#:~:text=Stopwords%20are%20words%20that%20are,Stopwords%20save%20storage%20space.","metadata":{}},{"cell_type":"code","source":"import en_core_sci_lg\n\n# medium model\nnlp = en_core_sci_lg.load(disable=[\"tagger\", \"parser\", \"ner\"])\nnlp.max_length = 2000000\n\n# New stop words list \ncustomize_stop_words = [\n    'what', 'a', 'and', 'are', 'as', 'at', 'be', 'author', 'by', 'for', 'if','into', 'in', 'is', 'it', 'no', 'not','of','on', 'or', 'such', 'that', 'the', 'their','then', 'there', 'this', 'to', 'was', 'will'\n   \n]\n\n# Mark them as stop words\nfor w in customize_stop_words:\n    nlp.vocab[w].is_stop = True","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:31:52.930796Z","iopub.execute_input":"2024-04-24T21:31:52.931297Z","iopub.status.idle":"2024-04-24T21:32:17.312065Z","shell.execute_reply.started":"2024-04-24T21:31:52.931261Z","shell.execute_reply":"2024-04-24T21:32:17.310857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def my_spacy_tokenizer(sentence):\n    # lowercase lemma, startchar and endchar of each word in sentence\n    return [(word.lemma_.lower(), word.idx, word.idx + len(word)) for word in nlp(sentence) \n            if not (word.like_num or word.is_stop or word.is_punct or word.is_space or len(word)==1)]","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:33:02.652629Z","iopub.execute_input":"2024-04-24T21:33:02.653223Z","iopub.status.idle":"2024-04-24T21:33:02.662448Z","shell.execute_reply.started":"2024-04-24T21:33:02.653173Z","shell.execute_reply":"2024-04-24T21:33:02.660493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SpacyTokenizer(Tokenizer):\n    \"\"\"\n    Customized tokenizer for Whoosh\n    \"\"\"\n    \n    def __init__(self, spacy_tokenizer):\n        self.spacy_tokenizer = spacy_tokenizer\n    \n    def __call__(self, value, positions = False, chars = False,\n                 keeporiginal = False, removestops = True,\n                 start_pos = 0, start_char = 0, mode = '', **kwargs):\n        \"\"\"\n        :param value: The unicode string to tokenize.\n        :param positions: Whether to record token positions in the token.\n        :param chars: Whether to record character offsets in the token.\n        :param start_pos: The position number of the first token. For example,\n            if you set start_pos=2, the tokens will be numbered 2,3,4,...\n            instead of 0,1,2,...\n        :param start_char: The offset of the first character of the first\n            token. For example, if you set start_char=2, the text \"aaa bbb\"\n            will have chars (2,5),(6,9) instead (0,3),(4,7).\n        \"\"\"\n        \n        assert isinstance(value, str), \"%r is not unicode\" % value\n        \n        spacy_tokens = self.spacy_tokenizer(value)\n        \n        t = Token(positions, chars, removestops=removestops, mode=mode)\n\n        for pos, spacy_token in enumerate(spacy_tokens):\n            t.text = spacy_token[0]\n            if keeporiginal:\n                t.original = t.text\n            t.stopped = False\n            if positions:\n                t.pos = start_pos + pos\n            if chars:\n                t.startchar = start_char + spacy_token[1]\n                t.endchar = start_char + spacy_token[2]\n            yield t","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:33:36.635041Z","iopub.execute_input":"2024-04-24T21:33:36.635675Z","iopub.status.idle":"2024-04-24T21:33:36.652176Z","shell.execute_reply.started":"2024-04-24T21:33:36.635631Z","shell.execute_reply":"2024-04-24T21:33:36.650589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#get hardcoded schema for the index\ndef get_search_schema(analyzer=StandardAnalyzer()):\n    schema = Schema(publication_number=ID(stored=True),\n                    title=TEXT(analyzer=analyzer, stored=True),\n                    abstract=TEXT(analyzer=analyzer, stored=True),\n                    claims=TEXT(analyzer=analyzer),\n                    description_text = TEXT(analyzer=analyzer)\n                   )\n    return schema\n\ndef add_documents_to_index(ix, df):\n    #create a writer object to add documents to the index\n    writer = ix.writer()\n\n    #now we can add documents to the index\n    for _, doc in df.iterrows():\n        writer.add_document(publication_number = str(doc.publication_number),\n                            title = str(doc.title),\n                            abstract = str(doc.abstract),\n                            claims = str(doc.claims),\n                            description = str(doc.description)\n#                             body_text = str(doc.body_text)                            \n                           )\n \n    writer.commit()\n\n    return\n\ndef create_search_index(search_schema):\n    if not os.path.exists('indexdir'):\n        os.mkdir('indexdir')\n        ix = index.create_in('indexdir', search_schema)\n        add_documents_to_index(ix, df)\n    else:           \n        #open an existing index object\n        ix = index.open_dir('indexdir')\n    return ix","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:43:30.963447Z","iopub.execute_input":"2024-04-24T21:43:30.964010Z","iopub.status.idle":"2024-04-24T21:43:30.977530Z","shell.execute_reply.started":"2024-04-24T21:43:30.963957Z","shell.execute_reply":"2024-04-24T21:43:30.975681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Create Search Index","metadata":{}},{"cell_type":"code","source":"my_analyzer = SpacyTokenizer(my_spacy_tokenizer)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:43:36.200220Z","iopub.execute_input":"2024-04-24T21:43:36.200776Z","iopub.status.idle":"2024-04-24T21:43:36.207771Z","shell.execute_reply.started":"2024-04-24T21:43:36.200738Z","shell.execute_reply":"2024-04-24T21:43:36.206056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"search_schema = get_search_schema(analyzer=my_analyzer)\nix = create_search_index(search_schema)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:43:40.326948Z","iopub.execute_input":"2024-04-24T21:43:40.327490Z","iopub.status.idle":"2024-04-24T21:43:40.336621Z","shell.execute_reply.started":"2024-04-24T21:43:40.327453Z","shell.execute_reply":"2024-04-24T21:43:40.334757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parser = MultifieldParser(['title', 'abstract'], schema=search_schema, group=OrGroup.factory(0.9))\nparser.add_plugin(SequencePlugin())","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:44:10.464328Z","iopub.execute_input":"2024-04-24T21:44:10.464789Z","iopub.status.idle":"2024-04-24T21:44:10.475954Z","shell.execute_reply.started":"2024-04-24T21:44:10.464753Z","shell.execute_reply":"2024-04-24T21:44:10.474277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Get Search Results","metadata":{}},{"cell_type":"code","source":"df.set_index('publication_number', inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T21:44:52.038741Z","iopub.execute_input":"2024-04-24T21:44:52.039187Z","iopub.status.idle":"2024-04-24T21:44:52.052101Z","shell.execute_reply.started":"2024-04-24T21:44:52.039155Z","shell.execute_reply":"2024-04-24T21:44:52.050612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Original code is_covid19=False  is_covid19 is one of the original columns","metadata":{}},{"cell_type":"code","source":"def search(query, lower=1975, upper=2013, kValue=5):\n    query_parsed = parser.parse(query)\n    with ix.searcher() as searcher:\n        results = searcher.search(query_parsed, limit = None)\n        output_dict = defaultdict(list)\n        for result in results:\n            output_dict['publication_number'].append(result['publication_number'])\n            output_dict['score'].append(result.score)\n        \n    search_results = pd.Series(output_dict['score'], index=pd.Index(output_dict['publication_number'], name='publication_number'))\n    search_results /= search_results.max()\n    \n    relevant_stuff = df.abstract.between(lower, upper)\n    #relevant_time = 2013\n\n    if only_claims:\n        temp = search_results[relevant_stuff & df.claims]\n\n    else:\n        temp = search_results[relevant_stuff]\n\n    if len(temp) == 0:\n        return -1\n\n    # Get top k matches\n    top_k = temp.nlargest(kValue)\n    \n    return top_k","metadata":{"execution":{"iopub.status.busy":"2024-04-24T22:36:23.409003Z","iopub.execute_input":"2024-04-24T22:36:23.409557Z","iopub.status.idle":"2024-04-24T22:36:23.422841Z","shell.execute_reply.started":"2024-04-24T22:36:23.409518Z","shell.execute_reply":"2024-04-24T22:36:23.421054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#I picked year 2013_12_parquet\n\npublication_number - Only patents published on or after 1975 were included in this column.\n\nIt was suppose to be only 2013, though that code was made for Covid and papers\n\nOriginal code : lower=1950, upper=2020 ","metadata":{}},{"cell_type":"code","source":"def searchQuery(query, attr=['publication_number', 'title', 'abstract', 'description'], kValue=3):\n    search_results = search(query, kValue)\n\n    if type(search_results) is int:\n        return []\n    #results = df.loc[search_results].reset_index()\n    results = pd.merge(search_results.to_frame(name='similarity'), df, on='publication_number', how='left').reset_index() \n\n    return results[attr].to_dict('records')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T22:37:17.235040Z","iopub.execute_input":"2024-04-24T22:37:17.235641Z","iopub.status.idle":"2024-04-24T22:37:17.244731Z","shell.execute_reply.started":"2024-04-24T22:37:17.235602Z","shell.execute_reply":"2024-04-24T22:37:17.242893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['title'][28895]","metadata":{"execution":{"iopub.status.busy":"2024-04-24T22:03:18.644220Z","iopub.execute_input":"2024-04-24T22:03:18.644871Z","iopub.status.idle":"2024-04-24T22:03:18.654521Z","shell.execute_reply.started":"2024-04-24T22:03:18.644827Z","shell.execute_reply":"2024-04-24T22:03:18.653199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#search('title: Methods, systems, and computer readable medium for providing video content over a network', kValue=3)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T22:43:52.383980Z","iopub.execute_input":"2024-04-24T22:43:52.384625Z","iopub.status.idle":"2024-04-24T22:43:52.390617Z","shell.execute_reply.started":"2024-04-24T22:43:52.384580Z","shell.execute_reply":"2024-04-24T22:43:52.389280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import whoosh\nfrom whoosh.fields import Schema, TEXT, ID\nfrom whoosh.index import create_in\nfrom whoosh.qparser import QueryParser\nimport os\n\n# Define schema\nschema = Schema(title=TEXT(phrase=True, stored=True), content=TEXT(phrase=True))\n\n# Create index directory\nif not os.path.exists(\"indexdir\"):\n    os.mkdir(\"indexdir\")\n\n# Create index\nix = create_in(\"indexdir\", schema)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T23:33:40.279911Z","iopub.execute_input":"2024-04-24T23:33:40.282985Z","iopub.status.idle":"2024-04-24T23:33:40.302202Z","shell.execute_reply.started":"2024-04-24T23:33:40.282890Z","shell.execute_reply":"2024-04-24T23:33:40.300401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Indexing documents\nwriter = ix.writer()\nwriter.add_document(title=u\"Methods, systems, and computer readable medium for providing video content over a network\", content=u\"This is the first document we've added!\")\nwriter.add_document(title=u\"Dynamically modifying the resources of a virtual server\", content=u\"The second one is even more interesting!\")\nwriter.commit()","metadata":{"execution":{"iopub.status.busy":"2024-04-24T23:38:17.679782Z","iopub.execute_input":"2024-04-24T23:38:17.680516Z","iopub.status.idle":"2024-04-24T23:38:17.704133Z","shell.execute_reply.started":"2024-04-24T23:38:17.680467Z","shell.execute_reply":"2024-04-24T23:38:17.702572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Searching\nwith ix.searcher() as searcher:\n    query = QueryParser(\"content\", ix.schema).parse(\"first\")\n    results = searcher.search(query)\n    for hit in results:\n        print(hit[\"title\"])","metadata":{"execution":{"iopub.status.busy":"2024-04-24T23:38:23.590529Z","iopub.execute_input":"2024-04-24T23:38:23.590987Z","iopub.status.idle":"2024-04-24T23:38:23.607987Z","shell.execute_reply.started":"2024-04-24T23:38:23.590955Z","shell.execute_reply":"2024-04-24T23:38:23.606229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with ix.searcher() as searcher:\n    query = QueryParser(\"content\", ix.schema).parse(\"computer\")\n    results = searcher.search(query)\n    for hit in results:\n        print(hit[\"title\"])","metadata":{"execution":{"iopub.status.busy":"2024-04-24T23:58:11.556430Z","iopub.execute_input":"2024-04-24T23:58:11.556940Z","iopub.status.idle":"2024-04-24T23:58:11.570767Z","shell.execute_reply.started":"2024-04-24T23:58:11.556906Z","shell.execute_reply":"2024-04-24T23:58:11.569212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#After \"Whooshing and FAILING for 5h:02m I gave it up. Even Sohier's Whoosh is module Not found\n\nEven forking it.","metadata":{}},{"cell_type":"code","source":"#By Sohier Dane https://www.kaggle.com/code/sohier/basic-whoosh-search-demo\n\nimport whoosh_utils","metadata":{"execution":{"iopub.status.busy":"2024-04-25T00:07:21.466128Z","iopub.execute_input":"2024-04-25T00:07:21.466607Z","iopub.status.idle":"2024-04-25T00:07:21.518695Z","shell.execute_reply.started":"2024-04-25T00:07:21.466574Z","shell.execute_reply":"2024-04-25T00:07:21.516386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Acknowledgements:\n\nSohier Dane https://www.kaggle.com/code/sohier/basic-whoosh-search-demo\n\nDaniel Wolffram https://www.kaggle.com/code/danielwolffram/whoosh-search\n\nDakinggg https://github.com/allenai/scispacy/blob/main/scripts/install_remote_packages.py\n\nMax Feinberg https://www.kaggle.com/code/mxfeinberg/using-whoosh-for-indexing-and-querying","metadata":{}}]}