{
  "id": 522359,
  "title": "13th place solution",
  "url": "/competitions/uspto-explainable-ai/writeups/camel-case-13th-place-solution",
  "author_name": "",
  "post_date": "2024-07-25T20:07:11.556198100Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to the organizers for the fun challenge!</p>\n<h3>Summary</h3>\n<ul>\n<li>Started in Python and got a public score of 0.53 using frequent pattern mining.</li>\n<li>Switched to C++, built my own searcher, created an ensemble of query generators, picking the best query for each test data row to get a public score of 0.76.</li>\n</ul>\n<h3>Original Python solution</h3>\n<p>I started out in Python, and used FP-Growth algorithm to find combinations of frequent CPC and title terms. I ranked these combinations by the number of targets they would match, and inversely by the estimated number of other non-target patents they would match. This resulted in a public score of 0.53.</p>\n<h3>Switching to C++</h3>\n<p>In experiments, I noticed that looking up the full data for random patents was extremely time-consuming, especially when that data would include descriptions. This seemed to be caused by the Parquet patent data files containing only a single row groups, making it required to read and decompress entire files even when only needing the data for a single patent.</p>\n<p>Around the same time I started thinking about the possibility of creating a search index and running queries against it in my submission notebook. This would make it possible to select the best query out of an array of generated queries for each row in the test dataset. However, generating a search index with Whoosh takes quite a lot of time.</p>\n<p>Based on those findings I came up with the following plan:</p>\n<ol>\n<li>Extract the compressed patent data from the Parquet files, tokenize their contents, and store the tokens to disk in a format optimized for my solution. This allow for fast random patent lookup, and fast lookup of a subset of a patent's fields (e.g. only retrieve the CPC and title tokens without parsing the tokens of other fields).</li>\n<li>At submission time, create an index containing patents that are most likely present in the test index. Assume all targets are present, and then gradually add popular neighbors of patents already in the index until the index contains 200k patents.  <br>\nThe index should store a bitset for each possible term denoting the patents that match it. Only consider <code>&lt;category&gt;:&lt;token&gt;</code> terms such as \"ti:method\" or \"detd:algorithm\" to limit the size of the index. Using such an index we can efficiently execute all queries using such terms with OR, AND, NOT, and XOR operators.  <br>\nThis format does not support proximity operators, wildcards, multi-word terms, and uncategorized terms. However, my earlier public score 0.53 submission only used single-word title terms and CPC terms with OR, AND, and XOR operators, so I was confident that the supported subset was enough to get a competitive leaderboard score.</li>\n<li>Create a bunch of query generators and use them to generate possible queries for each row in the test dataset, with per-query hyperparameter optimization where necessary. Run each generated query against the index created in step 2, and submit the one with the best result.</li>\n</ol>\n<p>Along with this plan I decided to switch to C++ for optimal performance. Kaggle notebooks run in an environment with GCC and CMake installed, so we can compile and run C++ code at submission time. Dependencies are pre-installed in a separate notebook and added through a private dataset.</p>\n<p>This plan mostly worked out, with some exceptions in step 2. Rather than 200k patents, I eventually switched to an index containing all 13M patents. I dropped support for description terms in this change to reduce the size of the index. I also added counts to the index, indicating how many times each patent matches each term. This resulted in a performance improvement in sorting patents that match the query to pick the top 50.</p>\n<h3>Reformatted patent data</h3>\n<p>I converted the original patent data files to my own format to allow for faster lookups. For each patent, my reformatted patent data contains unordered \"token -&gt; #occurrences\" mappings for the CPC codes, titles, abstracts, claims, and descriptions of all patents. This data is stored uncompressed in a custom binary format. I stored the mappings in a data file, and created an index file containing a separate mapping from publication number to the offset of its corresponding data in the data file. Together this format allows for fast lookup of random patent data.</p>\n<h3>Custom searcher</h3>\n<p>While building my own searcher I noticed several quirks:</p>\n<ul>\n<li>The XOR parser can be very slow. I've seen queries with just 10 XOR operators get stuck in the query parser in Python. Execution performance is not bad though, just need to make sure not to have too many XOR operators or submissions will time out.</li>\n<li><code>query_parser.parse(\"A OR B OR C\") == query_parser.parse(\"(A OR B) OR C\")</code> but <code>query_parser.parse(\"A XOR B XOR C\") != query_parser.parse(\"(A XOR B) XOR C\")</code>. <code>A XOR B XOR C</code> gets parsed to <code>((cpc:A OR cpc:B OR cpc:C) AND NOT (cpc:A AND cpc:B AND cpc:C))</code>, while <code>(A XOR B) XOR C</code> gets parsed to <code>((((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B)) OR cpc:C) AND NOT ((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B) AND cpc:C))</code>.</li>\n<li>If two patents match the query equally well, their order is determined by their Whoosh document id, which seems to be a number that increases every time a document is added to a Whoosh index. I couldn't figure out how to replicate this order in my own searcher.</li>\n</ul>\n<p>In the end my searcher was able to execute queries on an index containing 200k patents in less than 0.2ms on average with an in-memory index, and in less than 0.3ms on average with an index stored on disk (in a single thread on my own laptop).</p>\n<h3>Query generators</h3>\n<p>My final ensemble consisted out of four query generators:</p>\n<ol>\n<li>Single-term generator: this generator was used to test my implementation by generating queries containing a single term by picking the term with the maximum \"#targets matching the term / term selectivity\" score. Its queries didn't perform well, but as a testbed it was great.</li>\n<li>FP-Growth generator: this generator is essentially a clone of my original Python-based submission, but now running on abstract terms as well (for claims terms it is too slow).</li>\n<li>Optimizing generator: this generator supplies the most of my submitted queries. It attempts to find an optimal combination of target groups by starting with all targets in one group and then greedily swapping targets from one group to another (or to a new group). Target groups were converted to a query by creating a group of shared terms ranked by selectivity for each target group, and then equally dividing the number of available tokens across the groups. Terms inside each term group are then joined together by implicit AND operators, and groups of these AND-joined terms are joined together using OR and XOR operators (with a limit of 5 XOR operators in a query).</li>\n<li>Best-effort generator: this generator iteratively builds the query by constantly adding terms for the highest-priority target that isn't matched by the query yet. For example, if the first 3 targets match a partial query, this generator retrieves the terms of target 4 and adds the one with the lowest selectivity to the query. This generator mostly served as a fallback for cases where the optimizing generator fails to find a good query.</li>\n</ol>\n<h3>Local validation</h3>\n<p>I performed local validation on the first 2,500 rows of the neighbors data in <a href=\"https://www.kaggle.com/datasets/devinanzelmo/uspto-explainable-ai-validation-index\" target=\"_blank\">Devin Anzelmo's validation index</a>. I originally performed this local validation with an in-memory search index containing the 125k targets and 75k related patents, but switched to using a search index containing all 13M patents later as it gave more realistic results.</p>\n<h3>Code and data</h3>\n<ul>\n<li>Full source code: <a href=\"https://github.com/jmerle/uspto-explainable-ai\" target=\"_blank\">github.com/jmerle/uspto-explainable-ai</a></li>\n<li>Offline notebook (competition submission): <a href=\"https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble\" target=\"_blank\">kaggle.com/code/jmerle/uspto-explainable-ai-ensemble</a></li>\n<li>Online notebook (dependencies): <a href=\"https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies\" target=\"_blank\">kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies</a></li>\n<li>Reformatted competition data dataset: <a href=\"https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data\" target=\"_blank\">kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data</a></li>\n<li>Full search index dataset: <a href=\"https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index\" target=\"_blank\">kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index</a></li>\n<li>Original submission notebook (before switching to C++): <a href=\"https://www.kaggle.com/code/jmerle/uspto-explainable-ai\" target=\"_blank\">kaggle.com/code/jmerle/uspto-explainable-ai</a></li>\n</ul>",
  "messages": [
    {
      "id": "2936105",
      "postDate": "07/25/2024 20:07:11",
      "content": "<p>Thanks to the organizers for the fun challenge!</p>\n<h3>Summary</h3>\n<ul>\n<li>Started in Python and got a public score of 0.53 using frequent pattern mining.</li>\n<li>Switched to C++, built my own searcher, created an ensemble of query generators, picking the best query for each test data row to get a public score of 0.76.</li>\n</ul>\n<h3>Original Python solution</h3>\n<p>I started out in Python, and used FP-Growth algorithm to find combinations of frequent CPC and title terms. I ranked these combinations by the number of targets they would match, and inversely by the estimated number of other non-target patents they would match. This resulted in a public score of 0.53.</p>\n<h3>Switching to C++</h3>\n<p>In experiments, I noticed that looking up the full data for random patents was extremely time-consuming, especially when that data would include descriptions. This seemed to be caused by the Parquet patent data files containing only a single row groups, making it required to read and decompress entire files even when only needing the data for a single patent.</p>\n<p>Around the same time I started thinking about the possibility of creating a search index and running queries against it in my submission notebook. This would make it possible to select the best query out of an array of generated queries for each row in the test dataset. However, generating a search index with Whoosh takes quite a lot of time.</p>\n<p>Based on those findings I came up with the following plan:</p>\n<ol>\n<li>Extract the compressed patent data from the Parquet files, tokenize their contents, and store the tokens to disk in a format optimized for my solution. This allow for fast random patent lookup, and fast lookup of a subset of a patent's fields (e.g. only retrieve the CPC and title tokens without parsing the tokens of other fields).</li>\n<li>At submission time, create an index containing patents that are most likely present in the test index. Assume all targets are present, and then gradually add popular neighbors of patents already in the index until the index contains 200k patents.  <br>\nThe index should store a bitset for each possible term denoting the patents that match it. Only consider <code>&lt;category&gt;:&lt;token&gt;</code> terms such as \"ti:method\" or \"detd:algorithm\" to limit the size of the index. Using such an index we can efficiently execute all queries using such terms with OR, AND, NOT, and XOR operators.  <br>\nThis format does not support proximity operators, wildcards, multi-word terms, and uncategorized terms. However, my earlier public score 0.53 submission only used single-word title terms and CPC terms with OR, AND, and XOR operators, so I was confident that the supported subset was enough to get a competitive leaderboard score.</li>\n<li>Create a bunch of query generators and use them to generate possible queries for each row in the test dataset, with per-query hyperparameter optimization where necessary. Run each generated query against the index created in step 2, and submit the one with the best result.</li>\n</ol>\n<p>Along with this plan I decided to switch to C++ for optimal performance. Kaggle notebooks run in an environment with GCC and CMake installed, so we can compile and run C++ code at submission time. Dependencies are pre-installed in a separate notebook and added through a private dataset.</p>\n<p>This plan mostly worked out, with some exceptions in step 2. Rather than 200k patents, I eventually switched to an index containing all 13M patents. I dropped support for description terms in this change to reduce the size of the index. I also added counts to the index, indicating how many times each patent matches each term. This resulted in a performance improvement in sorting patents that match the query to pick the top 50.</p>\n<h3>Reformatted patent data</h3>\n<p>I converted the original patent data files to my own format to allow for faster lookups. For each patent, my reformatted patent data contains unordered \"token -&gt; #occurrences\" mappings for the CPC codes, titles, abstracts, claims, and descriptions of all patents. This data is stored uncompressed in a custom binary format. I stored the mappings in a data file, and created an index file containing a separate mapping from publication number to the offset of its corresponding data in the data file. Together this format allows for fast lookup of random patent data.</p>\n<h3>Custom searcher</h3>\n<p>While building my own searcher I noticed several quirks:</p>\n<ul>\n<li>The XOR parser can be very slow. I've seen queries with just 10 XOR operators get stuck in the query parser in Python. Execution performance is not bad though, just need to make sure not to have too many XOR operators or submissions will time out.</li>\n<li><code>query_parser.parse(\"A OR B OR C\") == query_parser.parse(\"(A OR B) OR C\")</code> but <code>query_parser.parse(\"A XOR B XOR C\") != query_parser.parse(\"(A XOR B) XOR C\")</code>. <code>A XOR B XOR C</code> gets parsed to <code>((cpc:A OR cpc:B OR cpc:C) AND NOT (cpc:A AND cpc:B AND cpc:C))</code>, while <code>(A XOR B) XOR C</code> gets parsed to <code>((((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B)) OR cpc:C) AND NOT ((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B) AND cpc:C))</code>.</li>\n<li>If two patents match the query equally well, their order is determined by their Whoosh document id, which seems to be a number that increases every time a document is added to a Whoosh index. I couldn't figure out how to replicate this order in my own searcher.</li>\n</ul>\n<p>In the end my searcher was able to execute queries on an index containing 200k patents in less than 0.2ms on average with an in-memory index, and in less than 0.3ms on average with an index stored on disk (in a single thread on my own laptop).</p>\n<h3>Query generators</h3>\n<p>My final ensemble consisted out of four query generators:</p>\n<ol>\n<li>Single-term generator: this generator was used to test my implementation by generating queries containing a single term by picking the term with the maximum \"#targets matching the term / term selectivity\" score. Its queries didn't perform well, but as a testbed it was great.</li>\n<li>FP-Growth generator: this generator is essentially a clone of my original Python-based submission, but now running on abstract terms as well (for claims terms it is too slow).</li>\n<li>Optimizing generator: this generator supplies the most of my submitted queries. It attempts to find an optimal combination of target groups by starting with all targets in one group and then greedily swapping targets from one group to another (or to a new group). Target groups were converted to a query by creating a group of shared terms ranked by selectivity for each target group, and then equally dividing the number of available tokens across the groups. Terms inside each term group are then joined together by implicit AND operators, and groups of these AND-joined terms are joined together using OR and XOR operators (with a limit of 5 XOR operators in a query).</li>\n<li>Best-effort generator: this generator iteratively builds the query by constantly adding terms for the highest-priority target that isn't matched by the query yet. For example, if the first 3 targets match a partial query, this generator retrieves the terms of target 4 and adds the one with the lowest selectivity to the query. This generator mostly served as a fallback for cases where the optimizing generator fails to find a good query.</li>\n</ol>\n<h3>Local validation</h3>\n<p>I performed local validation on the first 2,500 rows of the neighbors data in <a href=\"https://www.kaggle.com/datasets/devinanzelmo/uspto-explainable-ai-validation-index\" target=\"_blank\">Devin Anzelmo's validation index</a>. I originally performed this local validation with an in-memory search index containing the 125k targets and 75k related patents, but switched to using a search index containing all 13M patents later as it gave more realistic results.</p>\n<h3>Code and data</h3>\n<ul>\n<li>Full source code: <a href=\"https://github.com/jmerle/uspto-explainable-ai\" target=\"_blank\">github.com/jmerle/uspto-explainable-ai</a></li>\n<li>Offline notebook (competition submission): <a href=\"https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble\" target=\"_blank\">kaggle.com/code/jmerle/uspto-explainable-ai-ensemble</a></li>\n<li>Online notebook (dependencies): <a href=\"https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies\" target=\"_blank\">kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies</a></li>\n<li>Reformatted competition data dataset: <a href=\"https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data\" target=\"_blank\">kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data</a></li>\n<li>Full search index dataset: <a href=\"https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index\" target=\"_blank\">kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index</a></li>\n<li>Original submission notebook (before switching to C++): <a href=\"https://www.kaggle.com/code/jmerle/uspto-explainable-ai\" target=\"_blank\">kaggle.com/code/jmerle/uspto-explainable-ai</a></li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers for the fun challenge!\n\n### Summary\n- Started in Python and got a public score of 0.53 using frequent pattern mining.\n- Switched to C++, built my own searcher, created an ensemble of query generators, picking the best query for each test data row to get a public score of 0.76.\n\n### Original Python solution\nI started out in Python, and used FP-Growth algorithm to find combinations of frequent CPC and title terms. I ranked these combinations by the number of targets they would match, and inversely by the estimated number of other non-target patents they would match. This resulted in a public score of 0.53.\n\n### Switching to C++\nIn experiments, I noticed that looking up the full data for random patents was extremely time-consuming, especially when that data would include descriptions. This seemed to be caused by the Parquet patent data files containing only a single row groups, making it required to read and decompress entire files even when only needing the data for a single patent.\n\nAround the same time I started thinking about the possibility of creating a search index and running queries against it in my submission notebook. This would make it possible to select the best query out of an array of generated queries for each row in the test dataset. However, generating a search index with Whoosh takes quite a lot of time.\n\nBased on those findings I came up with the following plan:\n1. Extract the compressed patent data from the Parquet files, tokenize their contents, and store the tokens to disk in a format optimized for my solution. This allow for fast random patent lookup, and fast lookup of a subset of a patent's fields (e.g. only retrieve the CPC and title tokens without parsing the tokens of other fields).\n2. At submission time, create an index containing patents that are most likely present in the test index. Assume all targets are present, and then gradually add popular neighbors of patents already in the index until the index contains 200k patents.  \n   The index should store a bitset for each possible term denoting the patents that match it. Only consider `<category>:<token>` terms such as \"ti:method\" or \"detd:algorithm\" to limit the size of the index. Using such an index we can efficiently execute all queries using such terms with OR, AND, NOT, and XOR operators.  \n   This format does not support proximity operators, wildcards, multi-word terms, and uncategorized terms. However, my earlier public score 0.53 submission only used single-word title terms and CPC terms with OR, AND, and XOR operators, so I was confident that the supported subset was enough to get a competitive leaderboard score.\n3. Create a bunch of query generators and use them to generate possible queries for each row in the test dataset, with per-query hyperparameter optimization where necessary. Run each generated query against the index created in step 2, and submit the one with the best result.\n\nAlong with this plan I decided to switch to C++ for optimal performance. Kaggle notebooks run in an environment with GCC and CMake installed, so we can compile and run C++ code at submission time. Dependencies are pre-installed in a separate notebook and added through a private dataset.\n\nThis plan mostly worked out, with some exceptions in step 2. Rather than 200k patents, I eventually switched to an index containing all 13M patents. I dropped support for description terms in this change to reduce the size of the index. I also added counts to the index, indicating how many times each patent matches each term. This resulted in a performance improvement in sorting patents that match the query to pick the top 50.\n\n### Reformatted patent data\nI converted the original patent data files to my own format to allow for faster lookups. For each patent, my reformatted patent data contains unordered \"token -> #occurrences\" mappings for the CPC codes, titles, abstracts, claims, and descriptions of all patents. This data is stored uncompressed in a custom binary format. I stored the mappings in a data file, and created an index file containing a separate mapping from publication number to the offset of its corresponding data in the data file. Together this format allows for fast lookup of random patent data.\n\n### Custom searcher\nWhile building my own searcher I noticed several quirks:\n- The XOR parser can be very slow. I've seen queries with just 10 XOR operators get stuck in the query parser in Python. Execution performance is not bad though, just need to make sure not to have too many XOR operators or submissions will time out.\n- `query_parser.parse(\"A OR B OR C\") == query_parser.parse(\"(A OR B) OR C\")` but `query_parser.parse(\"A XOR B XOR C\") != query_parser.parse(\"(A XOR B) XOR C\")`. `A XOR B XOR C` gets parsed to `((cpc:A OR cpc:B OR cpc:C) AND NOT (cpc:A AND cpc:B AND cpc:C))`, while `(A XOR B) XOR C` gets parsed to `((((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B)) OR cpc:C) AND NOT ((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B) AND cpc:C))`.\n- If two patents match the query equally well, their order is determined by their Whoosh document id, which seems to be a number that increases every time a document is added to a Whoosh index. I couldn't figure out how to replicate this order in my own searcher.\n\nIn the end my searcher was able to execute queries on an index containing 200k patents in less than 0.2ms on average with an in-memory index, and in less than 0.3ms on average with an index stored on disk (in a single thread on my own laptop).\n\n### Query generators\nMy final ensemble consisted out of four query generators:\n1. Single-term generator: this generator was used to test my implementation by generating queries containing a single term by picking the term with the maximum \"#targets matching the term / term selectivity\" score. Its queries didn't perform well, but as a testbed it was great.\n2. FP-Growth generator: this generator is essentially a clone of my original Python-based submission, but now running on abstract terms as well (for claims terms it is too slow).\n3. Optimizing generator: this generator supplies the most of my submitted queries. It attempts to find an optimal combination of target groups by starting with all targets in one group and then greedily swapping targets from one group to another (or to a new group). Target groups were converted to a query by creating a group of shared terms ranked by selectivity for each target group, and then equally dividing the number of available tokens across the groups. Terms inside each term group are then joined together by implicit AND operators, and groups of these AND-joined terms are joined together using OR and XOR operators (with a limit of 5 XOR operators in a query).\n4. Best-effort generator: this generator iteratively builds the query by constantly adding terms for the highest-priority target that isn't matched by the query yet. For example, if the first 3 targets match a partial query, this generator retrieves the terms of target 4 and adds the one with the lowest selectivity to the query. This generator mostly served as a fallback for cases where the optimizing generator fails to find a good query.\n\n### Local validation\nI performed local validation on the first 2,500 rows of the neighbors data in [Devin Anzelmo's validation index](https://www.kaggle.com/datasets/devinanzelmo/uspto-explainable-ai-validation-index). I originally performed this local validation with an in-memory search index containing the 125k targets and 75k related patents, but switched to using a search index containing all 13M patents later as it gave more realistic results.\n\n### Code and data\n- Full source code: [github.com/jmerle/uspto-explainable-ai](https://github.com/jmerle/uspto-explainable-ai)\n- Offline notebook (competition submission): [kaggle.com/code/jmerle/uspto-explainable-ai-ensemble](https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble)\n- Online notebook (dependencies): [kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies](https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies)\n- Reformatted competition data dataset: [kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data](https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data)\n- Full search index dataset: [kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index](https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index)\n- Original submission notebook (before switching to C++): [kaggle.com/code/jmerle/uspto-explainable-ai](https://www.kaggle.com/code/jmerle/uspto-explainable-ai)",
      "votes": null
    },
    {
      "id": "2936282",
      "postDate": "07/26/2024 03:21:15",
      "content": "<p>Congrats on your medal, <a href=\"https://www.kaggle.com/jmerle\" target=\"_blank\">@jmerle</a>, and thanks for your complete remarks about your work. I’m curious—how much time did you save by switching to C++, and have you considered using C? I’d love to hear your thoughts!</p>",
      "rawMarkdown": "Congrats on your medal, @jmerle, and thanks for your complete remarks about your work. I’m curious—how much time did you save by switching to C++, and have you considered using C? I’d love to hear your thoughts!",
      "votes": null
    },
    {
      "id": "2937197",
      "postDate": "07/26/2024 19:43:10",
      "content": "<p>In terms of development time it was probably a bit slower than doing it in Python. However, in terms of execution time I expect my C++ code to be much faster than its equivalent in Python would be, although I never tested that.</p>\n<p>I never considered C because I prefer C++ when I have the choice. I'm more used to it, it has more language features available, and it has many more libraries available. I did use a library written in C in my submission, which is also one of the things I like about C++ (the ability to use libraries written in both C and in C++).</p>",
      "rawMarkdown": "In terms of development time it was probably a bit slower than doing it in Python. However, in terms of execution time I expect my C++ code to be much faster than its equivalent in Python would be, although I never tested that.\n\nI never considered C because I prefer C++ when I have the choice. I'm more used to it, it has more language features available, and it has many more libraries available. I did use a library written in C in my submission, which is also one of the things I like about C++ (the ability to use libraries written in both C and in C++).",
      "votes": null
    },
    {
      "id": "2937684",
      "postDate": "07/27/2024 10:39:54",
      "content": "<p><a href=\"https://www.kaggle.com/jmerle\" target=\"_blank\">@jmerle</a> Thank you for sharing your approach. I read through your GitHub-you really did comprehensive work. You should feel proud of yourself!</p>",
      "rawMarkdown": "jmerle Thank you for sharing your approach. I read through your GitHub-you really did comprehensive work. You should feel proud of yourself!",
      "votes": null
    },
    {
      "id": "2938674",
      "postDate": "07/28/2024 11:44:32",
      "content": "<p>Msn thats a great insight ! Thankyou for this</p>",
      "rawMarkdown": "Msn thats a great insight ! Thankyou for this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2936282,
      "author_name": "gowillgo",
      "author_url": "",
      "post_date": "07/26/2024 03:21:15",
      "content": "<p>Congrats on your medal, <a href=\"https://www.kaggle.com/jmerle\" target=\"_blank\">@jmerle</a>, and thanks for your complete remarks about your work. I’m curious—how much time did you save by switching to C++, and have you considered using C? I’d love to hear your thoughts!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2937197,
          "author_name": "jmerle",
          "author_url": "",
          "post_date": "07/26/2024 19:43:10",
          "content": "<p>In terms of development time it was probably a bit slower than doing it in Python. However, in terms of execution time I expect my C++ code to be much faster than its equivalent in Python would be, although I never tested that.</p>\n<p>I never considered C because I prefer C++ when I have the choice. I'm more used to it, it has more language features available, and it has many more libraries available. I did use a library written in C in my submission, which is also one of the things I like about C++ (the ability to use libraries written in both C and in C++).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2937684,
              "author_name": "gowillgo",
              "author_url": "",
              "post_date": "07/27/2024 10:39:54",
              "content": "<p><a href=\"https://www.kaggle.com/jmerle\" target=\"_blank\">@jmerle</a> Thank you for sharing your approach. I read through your GitHub-you really did comprehensive work. You should feel proud of yourself!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2938674,
      "author_name": "aadiar",
      "author_url": "",
      "post_date": "07/28/2024 11:44:32",
      "content": "<p>Msn thats a great insight ! Thankyou for this</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2936105": "Thanks to the organizers for the fun challenge!\n\n### Summary\n- Started in Python and got a public score of 0.53 using frequent pattern mining.\n- Switched to C++, built my own searcher, created an ensemble of query generators, picking the best query for each test data row to get a public score of 0.76.\n\n### Original Python solution\nI started out in Python, and used FP-Growth algorithm to find combinations of frequent CPC and title terms. I ranked these combinations by the number of targets they would match, and inversely by the estimated number of other non-target patents they would match. This resulted in a public score of 0.53.\n\n### Switching to C++\nIn experiments, I noticed that looking up the full data for random patents was extremely time-consuming, especially when that data would include descriptions. This seemed to be caused by the Parquet patent data files containing only a single row groups, making it required to read and decompress entire files even when only needing the data for a single patent.\n\nAround the same time I started thinking about the possibility of creating a search index and running queries against it in my submission notebook. This would make it possible to select the best query out of an array of generated queries for each row in the test dataset. However, generating a search index with Whoosh takes quite a lot of time.\n\nBased on those findings I came up with the following plan:\n1. Extract the compressed patent data from the Parquet files, tokenize their contents, and store the tokens to disk in a format optimized for my solution. This allow for fast random patent lookup, and fast lookup of a subset of a patent's fields (e.g. only retrieve the CPC and title tokens without parsing the tokens of other fields).\n2. At submission time, create an index containing patents that are most likely present in the test index. Assume all targets are present, and then gradually add popular neighbors of patents already in the index until the index contains 200k patents.  \n   The index should store a bitset for each possible term denoting the patents that match it. Only consider `<category>:<token>` terms such as \"ti:method\" or \"detd:algorithm\" to limit the size of the index. Using such an index we can efficiently execute all queries using such terms with OR, AND, NOT, and XOR operators.  \n   This format does not support proximity operators, wildcards, multi-word terms, and uncategorized terms. However, my earlier public score 0.53 submission only used single-word title terms and CPC terms with OR, AND, and XOR operators, so I was confident that the supported subset was enough to get a competitive leaderboard score.\n3. Create a bunch of query generators and use them to generate possible queries for each row in the test dataset, with per-query hyperparameter optimization where necessary. Run each generated query against the index created in step 2, and submit the one with the best result.\n\nAlong with this plan I decided to switch to C++ for optimal performance. Kaggle notebooks run in an environment with GCC and CMake installed, so we can compile and run C++ code at submission time. Dependencies are pre-installed in a separate notebook and added through a private dataset.\n\nThis plan mostly worked out, with some exceptions in step 2. Rather than 200k patents, I eventually switched to an index containing all 13M patents. I dropped support for description terms in this change to reduce the size of the index. I also added counts to the index, indicating how many times each patent matches each term. This resulted in a performance improvement in sorting patents that match the query to pick the top 50.\n\n### Reformatted patent data\nI converted the original patent data files to my own format to allow for faster lookups. For each patent, my reformatted patent data contains unordered \"token -> #occurrences\" mappings for the CPC codes, titles, abstracts, claims, and descriptions of all patents. This data is stored uncompressed in a custom binary format. I stored the mappings in a data file, and created an index file containing a separate mapping from publication number to the offset of its corresponding data in the data file. Together this format allows for fast lookup of random patent data.\n\n### Custom searcher\nWhile building my own searcher I noticed several quirks:\n- The XOR parser can be very slow. I've seen queries with just 10 XOR operators get stuck in the query parser in Python. Execution performance is not bad though, just need to make sure not to have too many XOR operators or submissions will time out.\n- `query_parser.parse(\"A OR B OR C\") == query_parser.parse(\"(A OR B) OR C\")` but `query_parser.parse(\"A XOR B XOR C\") != query_parser.parse(\"(A XOR B) XOR C\")`. `A XOR B XOR C` gets parsed to `((cpc:A OR cpc:B OR cpc:C) AND NOT (cpc:A AND cpc:B AND cpc:C))`, while `(A XOR B) XOR C` gets parsed to `((((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B)) OR cpc:C) AND NOT ((cpc:A OR cpc:B) AND NOT (cpc:A AND cpc:B) AND cpc:C))`.\n- If two patents match the query equally well, their order is determined by their Whoosh document id, which seems to be a number that increases every time a document is added to a Whoosh index. I couldn't figure out how to replicate this order in my own searcher.\n\nIn the end my searcher was able to execute queries on an index containing 200k patents in less than 0.2ms on average with an in-memory index, and in less than 0.3ms on average with an index stored on disk (in a single thread on my own laptop).\n\n### Query generators\nMy final ensemble consisted out of four query generators:\n1. Single-term generator: this generator was used to test my implementation by generating queries containing a single term by picking the term with the maximum \"#targets matching the term / term selectivity\" score. Its queries didn't perform well, but as a testbed it was great.\n2. FP-Growth generator: this generator is essentially a clone of my original Python-based submission, but now running on abstract terms as well (for claims terms it is too slow).\n3. Optimizing generator: this generator supplies the most of my submitted queries. It attempts to find an optimal combination of target groups by starting with all targets in one group and then greedily swapping targets from one group to another (or to a new group). Target groups were converted to a query by creating a group of shared terms ranked by selectivity for each target group, and then equally dividing the number of available tokens across the groups. Terms inside each term group are then joined together by implicit AND operators, and groups of these AND-joined terms are joined together using OR and XOR operators (with a limit of 5 XOR operators in a query).\n4. Best-effort generator: this generator iteratively builds the query by constantly adding terms for the highest-priority target that isn't matched by the query yet. For example, if the first 3 targets match a partial query, this generator retrieves the terms of target 4 and adds the one with the lowest selectivity to the query. This generator mostly served as a fallback for cases where the optimizing generator fails to find a good query.\n\n### Local validation\nI performed local validation on the first 2,500 rows of the neighbors data in [Devin Anzelmo's validation index](https://www.kaggle.com/datasets/devinanzelmo/uspto-explainable-ai-validation-index). I originally performed this local validation with an in-memory search index containing the 125k targets and 75k related patents, but switched to using a search index containing all 13M patents later as it gave more realistic results.\n\n### Code and data\n- Full source code: [github.com/jmerle/uspto-explainable-ai](https://github.com/jmerle/uspto-explainable-ai)\n- Offline notebook (competition submission): [kaggle.com/code/jmerle/uspto-explainable-ai-ensemble](https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble)\n- Online notebook (dependencies): [kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies](https://www.kaggle.com/code/jmerle/uspto-explainable-ai-ensemble-dependencies)\n- Reformatted competition data dataset: [kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data](https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-reformatted-patent-data)\n- Full search index dataset: [kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index](https://www.kaggle.com/datasets/jmerle/uspto-explainable-ai-full-search-index)\n- Original submission notebook (before switching to C++): [kaggle.com/code/jmerle/uspto-explainable-ai](https://www.kaggle.com/code/jmerle/uspto-explainable-ai)",
    "2936282": "Congrats on your medal, @jmerle, and thanks for your complete remarks about your work. I’m curious—how much time did you save by switching to C++, and have you considered using C? I’d love to hear your thoughts!",
    "2937197": "In terms of development time it was probably a bit slower than doing it in Python. However, in terms of execution time I expect my C++ code to be much faster than its equivalent in Python would be, although I never tested that.\n\nI never considered C because I prefer C++ when I have the choice. I'm more used to it, it has more language features available, and it has many more libraries available. I did use a library written in C in my submission, which is also one of the things I like about C++ (the ability to use libraries written in both C and in C++).",
    "2937684": "jmerle Thank you for sharing your approach. I read through your GitHub-you really did comprehensive work. You should feel proud of yourself!",
    "2938674": "Msn thats a great insight ! Thankyou for this"
  },
  "source": "meta"
}