{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59575,"databundleVersionId":8060720,"sourceType":"competition"},{"sourceId":11371,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":5171}],"dockerImageVersionId":30699,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on April 28, 2024. By Marília Prata, mpwolke","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-28T22:40:38.013572Z","iopub.execute_input":"2024-04-28T22:40:38.013928Z","iopub.status.idle":"2024-04-28T22:40:40.769402Z","shell.execute_reply.started":"2024-04-28T22:40:38.013889Z","shell.execute_reply":"2024-04-28T22:40:40.768452Z"},"_kg_hide-output":true,"_kg_hide-input":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![](https://image.slidesharecdn.com/patent-170125095602/85/patent-search-14-320.jpg)https://pt.slideshare.net/Brainleague/patent-search-71367455","metadata":{}},{"cell_type":"markdown","source":"#USPTO Default Operators: OR, AND, ADJ, NEAR, SAME, WITH.\n\nhttps://ppubs.uspto.gov/pubwebapp/","metadata":{}},{"cell_type":"markdown","source":"#To Reduce the number of Parquet file rows\n\nI reduced cause Epochs with 28899 rows took 50 minutes Just to ten percent of the data. Therefore, the Epoch (Gemma_lm.fit) would take at least 500 minutes and I don't even had arrived on the boolean Queries to check any result and leave any comment.","metadata":{}},{"cell_type":"code","source":"#By StackOverflow https://stackoverflow.com/questions/53982871/pandas-reading-first-n-rows-from-parquet-file\n\nfrom pyarrow.parquet import ParquetFile\nimport pyarrow as pa \n\npf = ParquetFile('../input/uspto-explainable-ai/patent_data/2013_12.parquet') \nfirst_thousand_rows = next(pf.iter_batches(batch_size = 1000)) \ndf = pa.Table.from_batches([first_thousand_rows]).to_pandas()","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:41:20.912432Z","iopub.execute_input":"2024-04-28T22:41:20.913234Z","iopub.status.idle":"2024-04-28T22:41:24.479052Z","shell.execute_reply.started":"2024-04-28T22:41:20.913200Z","shell.execute_reply":"2024-04-28T22:41:24.478103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.tail()","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:41:28.979542Z","iopub.execute_input":"2024-04-28T22:41:28.980142Z","iopub.status.idle":"2024-04-28T22:41:28.997610Z","shell.execute_reply.started":"2024-04-28T22:41:28.980112Z","shell.execute_reply":"2024-04-28T22:41:28.996717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install whoosh","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-28T22:41:37.723983Z","iopub.execute_input":"2024-04-28T22:41:37.724604Z","iopub.status.idle":"2024-04-28T22:41:51.352615Z","shell.execute_reply.started":"2024-04-28T22:41:37.724572Z","shell.execute_reply":"2024-04-28T22:41:51.351484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By SeshuRajup https://www.kaggle.com/code/seshurajup/patent-professionals-uspto-whoosh-utils-eda/notebook\n\nimport re\nimport whoosh\nimport whoosh.analysis\n\nNUMBER_REGEX = re.compile(r'^(\\d+|\\d{1,3}(,\\d{3})*)(\\.\\d+)?$')\n\nclass NumberFilter(whoosh.analysis.Filter):\n    def __call__(self, tokens):\n        for t in tokens:\n            if not NUMBER_REGEX.match(t.text):\n                yield t\n                \nBRS_STOPWORDS = ['an', 'are', 'by', 'for', 'if', 'into', 'is', 'no', 'not', 'of', 'on', 'such',\n        'that', 'the', 'their', 'then', 'there', 'these', 'they', 'this', 'to', 'was', 'will']\ntext_analyzer = whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:41:56.493374Z","iopub.execute_input":"2024-04-28T22:41:56.493745Z","iopub.status.idle":"2024-04-28T22:41:56.514436Z","shell.execute_reply.started":"2024-04-28T22:41:56.493713Z","shell.execute_reply":"2024-04-28T22:41:56.513631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#X-Ray tube Patent","metadata":{}},{"cell_type":"code","source":"#By SeshuRajup https://www.kaggle.com/code/seshurajup/patent-professionals-uspto-whoosh-utils-eda/notebook\n\nfrom IPython.display import display, HTML\ndef parse_data(section, text):\n    display(HTML(f\"<b style='color:blue'>{section.title()}</b>: <span style='color:blue'>{text}<span>\"))\n    tokens = text_analyzer(text)\n    display(HTML(f\"<b style='color:green'>{section.title()} Tokens</b>: <span style='color:green'>{str([token.text for token in tokens])}</span>\"))\n    \nfor i in range(len(df['publication_number'])):\n    if i > 0:  #Original was higher number\n        break\n    \n    pubnum = df['publication_number'][998]\n    title = df['title'][998]\n    abstract = df['abstract'][998]\n    #claims = patent_details['claims'][998]\n    #description = patent_details['description'][998]\n    display(HTML(f\"<br/>Publication number: <b style='color:red'>{pubnum}</b>\"))\n    parse_data('title', title)\n    parse_data('abstract', abstract)\n    #parse_data('claims', claims)\n    #parse_data('description', description)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-28T22:42:02.020935Z","iopub.execute_input":"2024-04-28T22:42:02.021302Z","iopub.status.idle":"2024-04-28T22:42:02.038092Z","shell.execute_reply.started":"2024-04-28T22:42:02.021273Z","shell.execute_reply":"2024-04-28T22:42:02.037204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Solar Module Patent","metadata":{}},{"cell_type":"code","source":"#By SeshuRajup https://www.kaggle.com/code/seshurajup/patent-professionals-uspto-whoosh-utils-eda/notebook\n\nfor i in range(len(df['publication_number'])):\n    if i > 0:  #Original was higher number\n        break\n    \n    pubnum = df['publication_number'][250]\n    title = df['title'][250]\n    abstract = df['abstract'][250]\n    #claims = patent_details['claims'][250]\n    #description = patent_details['description'][250]\n    display(HTML(f\"<br/>Publication number: <b style='color:red'>{pubnum}</b>\"))\n    parse_data('title', title)\n    parse_data('abstract', abstract)\n    #parse_data('claims', claims)\n    #parse_data('description', description)","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:42:07.710452Z","iopub.execute_input":"2024-04-28T22:42:07.711192Z","iopub.status.idle":"2024-04-28T22:42:07.724881Z","shell.execute_reply.started":"2024-04-28T22:42:07.711148Z","shell.execute_reply":"2024-04-28T22:42:07.723944Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Install Keras 3 last. See https://keras.io/getting_started/ for more details.\n!pip install -q -U keras-nlp\n!pip install -q -U keras>=3\n\nimport os\n\nos.environ[\"KERAS_BACKEND\"] = \"jax\"  # Or \"torch\" or \"tensorflow\".\n# Avoid memory fragmentation on JAX backend.\nos.environ[\"XLA_PYTHON_CLIENT_MEM_FRACTION\"]=\"1.00\"\n\nimport keras\nimport keras_nlp","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-28T22:42:12.949581Z","iopub.execute_input":"2024-04-28T22:42:12.950012Z","iopub.status.idle":"2024-04-28T22:42:54.469114Z","shell.execute_reply.started":"2024-04-28T22:42:12.949977Z","shell.execute_reply":"2024-04-28T22:42:54.468094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gemma_lm = keras_nlp.models.GemmaCausalLM.from_preset(\"gemma_2b_en\")\n#gemma_lm.summary()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-28T22:43:14.917759Z","iopub.execute_input":"2024-04-28T22:43:14.918748Z","iopub.status.idle":"2024-04-28T22:44:15.674729Z","shell.execute_reply.started":"2024-04-28T22:43:14.918710Z","shell.execute_reply":"2024-04-28T22:44:15.673858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By York Yong https://www.kaggle.com/code/yorkyong/gemma-trial-ace-a-data-science-interview\n\nuspto_dataset = []\n    \nfor index, row in df.iterrows():\n    question, answer = row['abstract'], row['title']\n    template = (f\"abstract:\\n{question}\\n\\ntitle:\\n{answer}\")\n    uspto_dataset.append(template)","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:44:22.496546Z","iopub.execute_input":"2024-04-28T22:44:22.496908Z","iopub.status.idle":"2024-04-28T22:44:22.570995Z","shell.execute_reply.started":"2024-04-28T22:44:22.496881Z","shell.execute_reply":"2024-04-28T22:44:22.570023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#LoRA","metadata":{}},{"cell_type":"code","source":"# Enable LoRA for the model and set the LoRA rank to 64.\ngemma_lm.backbone.enable_lora(rank=64)","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:44:29.835819Z","iopub.execute_input":"2024-04-28T22:44:29.836190Z","iopub.status.idle":"2024-04-28T22:44:30.239388Z","shell.execute_reply.started":"2024-04-28T22:44:29.836150Z","shell.execute_reply":"2024-04-28T22:44:30.238417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Control Memory Usage","metadata":{}},{"cell_type":"code","source":"# Limit the input sequence length to 512 (to control memory usage).\ngemma_lm.preprocessor.sequence_length = 512\n# Use AdamW (a common optimizer for transformer models).\noptimizer = keras.optimizers.AdamW(\n    learning_rate=5e-5,\n    weight_decay=0.01,\n)\n# Exclude layernorm and bias terms from decay.\noptimizer.exclude_from_weight_decay(var_names=[\"bias\", \"scale\"])\n\ngemma_lm.compile(\n    loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),\n    optimizer=optimizer,\n    weighted_metrics=[keras.metrics.SparseCategoricalAccuracy()],\n)","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:44:41.879992Z","iopub.execute_input":"2024-04-28T22:44:41.880381Z","iopub.status.idle":"2024-04-28T22:44:41.891811Z","shell.execute_reply.started":"2024-04-28T22:44:41.880354Z","shell.execute_reply":"2024-04-28T22:44:41.891034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Epochs started 19:44 till 19:57","metadata":{}},{"cell_type":"code","source":"%%time\n\ngemma_lm.fit(uspto_dataset, epochs=1, batch_size=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:44:49.714711Z","iopub.execute_input":"2024-04-28T22:44:49.715074Z","iopub.status.idle":"2024-04-28T22:57:27.362362Z","shell.execute_reply.started":"2024-04-28T22:44:49.715045Z","shell.execute_reply":"2024-04-28T22:57:27.361407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Boolean search operator \"OR\"","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"Were x-ray OR Solar Modules US Patents published?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:58:06.488043Z","iopub.execute_input":"2024-04-28T22:58:06.488984Z","iopub.status.idle":"2024-04-28T22:58:25.665361Z","shell.execute_reply.started":"2024-04-28T22:58:06.488945Z","shell.execute_reply":"2024-04-28T22:58:25.664380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Boolean search operator AND","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X-Ray AND Solar Module Patents were published?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:58:37.277355Z","iopub.execute_input":"2024-04-28T22:58:37.277726Z","iopub.status.idle":"2024-04-28T22:58:44.202123Z","shell.execute_reply.started":"2024-04-28T22:58:37.277698Z","shell.execute_reply":"2024-04-28T22:58:44.201105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#I was dealing with a 2013_12 parquet file (December 2013)","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X_Ray Patents + Solar Module Patents were published?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:58:50.618455Z","iopub.execute_input":"2024-04-28T22:58:50.618872Z","iopub.status.idle":"2024-04-28T22:58:57.516432Z","shell.execute_reply.started":"2024-04-28T22:58:50.618840Z","shell.execute_reply":"2024-04-28T22:58:57.515548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#I've just replaced the Operator AND  to +. \n\nThe model repeated \"I am not sure if this is a good thing or not.\"","metadata":{}},{"cell_type":"markdown","source":"#Operator \"NEAR\"","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X_Ray Patents NEAR Anode Unit were published??\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:59:07.506264Z","iopub.execute_input":"2024-04-28T22:59:07.506625Z","iopub.status.idle":"2024-04-28T22:59:14.416891Z","shell.execute_reply.started":"2024-04-28T22:59:07.506596Z","shell.execute_reply":"2024-04-28T22:59:14.415899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Operator \"NOT\"","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X_Ray Patents NOT Publication number: US-2013322602-A1?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:59:24.262497Z","iopub.execute_input":"2024-04-28T22:59:24.263470Z","iopub.status.idle":"2024-04-28T22:59:30.802513Z","shell.execute_reply.started":"2024-04-28T22:59:24.263427Z","shell.execute_reply":"2024-04-28T22:59:30.801376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X-Ray Patents - US Patent Publication number: US-2013322601-A1?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:59:38.865200Z","iopub.execute_input":"2024-04-28T22:59:38.865544Z","iopub.status.idle":"2024-04-28T22:59:45.345990Z","shell.execute_reply.started":"2024-04-28T22:59:38.865520Z","shell.execute_reply":"2024-04-28T22:59:45.344999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"Solar Module Patents - US Patent Publication number: US-2013319518-A12?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T22:59:51.286886Z","iopub.execute_input":"2024-04-28T22:59:51.287275Z","iopub.status.idle":"2024-04-28T22:59:57.782991Z","shell.execute_reply.started":"2024-04-28T22:59:51.287246Z","shell.execute_reply":"2024-04-28T22:59:57.781976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Operator \"ADJ\" (Adjacente)","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X-Ray Patent ADJ Publication number: US-2013322602-A1?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T23:00:01.643228Z","iopub.execute_input":"2024-04-28T23:00:01.643867Z","iopub.status.idle":"2024-04-28T23:00:08.151068Z","shell.execute_reply.started":"2024-04-28T23:00:01.643839Z","shell.execute_reply":"2024-04-28T23:00:08.150095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Operator \"Same\"","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X-Ray Patent Publication number: US-2013322602-A1 SAME title?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T23:00:11.399222Z","iopub.execute_input":"2024-04-28T23:00:11.399601Z","iopub.status.idle":"2024-04-28T23:00:17.906123Z","shell.execute_reply.started":"2024-04-28T23:00:11.399572Z","shell.execute_reply":"2024-04-28T23:00:17.905222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Operator 'With'","metadata":{}},{"cell_type":"code","source":"%%time\nprint(gemma_lm.generate(\"X-Ray Patent WITH Publication number: US-2013322602-A1?\", max_length=256))","metadata":{"execution":{"iopub.status.busy":"2024-04-28T23:00:27.119676Z","iopub.execute_input":"2024-04-28T23:00:27.120045Z","iopub.status.idle":"2024-04-28T23:00:33.658746Z","shell.execute_reply.started":"2024-04-28T23:00:27.120016Z","shell.execute_reply":"2024-04-28T23:00:33.657662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Maybe I should propose better queries\n\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcRIjClknXgfpLFG_bOwftSkTWNV5qpvflgY6w&s)ResearchGate","metadata":{}},{"cell_type":"markdown","source":"#AI’s Future Role in Patent Search\n\n\"While AI holds great promise in refining patent searches, it’s crucial to view it as a complementary tool rather than a replacement for human intellect. Challenges persist, such as understanding the context of a Person Having Ordinary Skill in the Art (POSITA), managing complex situations, and addressing biases in AI. Despite its potential, AI has yet to reach the level of insight that a human examiner brings to the process.\"\n\nhttps://www.legaladvantage.net/blog/unlocking-ais-impact-on-uspto-patent-search/","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nYork Yong https://www.kaggle.com/code/yorkyong/gemma-trial-ace-a-data-science-interview\n\nSeshuRajup https://www.kaggle.com/code/seshurajup/patent-professionals-uspto-whoosh-utils-eda/notebook\n\nmpwolke https://www.kaggle.com/code/mpwolke/what-is-1-1-gemma","metadata":{}}]}