{
  "id": 522301,
  "title": "12th place - Disjunctions of bigrams, trigrams only",
  "url": "/competitions/uspto-explainable-ai/discussion/522301",
  "author_name": "dt",
  "post_date": "2024-07-25T13:01:24.982000",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My initial solution was unigram combinations of all fields, but after discovering the ‘magic character’* in phrase search and getting good initial results, I pivoted to using bigrams and trigrams only. Ultimately it wasn’t a good decision for the scoring and my only consolation is that the majority of bigrams/trigrams query terms in my final solution are quite meaningful. It’s also quite cool when target patents can be matched by a single bigram/trigram.</p>\n<p>I’ll go into detail about 3 main parts of my final solution:</p>\n<p><a href=\"#preprocessing\">Preprocessing</a><br>\n<a href=\"#weighted-set-cover-algorithm\">Weighted set cover algorithm</a><br>\n<a href=\"#parameter-optimization\">Parameter optimization</a></p>\n<p>*Side note:<br>\nThe ‘magic character’ that I’m referring to is the usage of special characters to join words in phrases, which is the only trick I used in my solution. The character I used to join phrases was ‘-’. The token savings of this trick in my solution was actually not that significant in practice, due to the following reasons:</p>\n<ul>\n<li>Because I want to use meaningful bigrams and trigrams in all fields, it is more important to find ngrams that can target span patent groups. In a significant proportion of patent groups I think there actually aren’t such bigrams/trigrams, so it imposes an upper bound on the possible cover with 50 tokens. In some patent groups however it does work very nicely and you can actually find a bigram/trigram that can cover all 50 neighbors without false positives (in such cases the trick also doesn’t matter since we only need a few tokens)</li>\n<li>Phrase search in description has VERY expensive query time and a major issue I faced in the last 2 weeks after I created my final phrase dataset. I came to realise that some bigrams/trigrams can even take more than 20s on their own and I had to implement multiple changes to get my notebook to score reliably. These findings are documented in a separate post <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522303\" target=\"_blank\">here</a></li>\n<li>Phrase search in Whoosh had some inconsistencies such as the one described <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/519755#2919472\" target=\"_blank\">here</a>, that is complex to account for in the preprocessing pipeline</li>\n</ul>\n<p>In summary, phrase search in title, abstract and claims is quite stable, but in description field … <em>here be dragons</em>. At the same time description field accounts for majority of the query results for this approach and having only title, claims and abstract would score very poorly.</p>\n<h2>Preprocessing</h2>\n<p>The goal of the preprocessing pipeline is to generate phrase ranks for all patents such that for each target patent, we can fetch ranked phrases for each neighbour and assemble them together quickly. A finite state automaton built using Aho-Corasick is the key part of this pipeline.</p>\n<h6>SpaCy NLP + CountVectorizer</h6>\n<p>To extract noun chunks I used SpaCy’s nlp pipeline inside a batched CountVectorizer with optimizations like</p>\n<ul>\n<li>disabling lemmatizer and ner in nlp pipeline</li>\n<li>analyze only first 1 million characters in description text (SpaCy’s default maximum length for nlp pipeline is 1 million)</li>\n<li>binary True and max_df of 0.1 for CountVectorizer</li>\n<li>truncating or dropping noun chunks with special characters</li>\n<li>dropping noun chunks with count &gt; 150 after combining batches</li>\n<li>clean the combined vocabulary, check for length, Shannon entropy, etc.  </li>\n</ul>\n<p>Final output is a vocabulary of phrases (the count obtained here is an estimate, the true count will be obtained later through the fsa)</p>\n<h6>Build fsa from vocabulary</h6>\n<p>A finite state automaton using Aho-Corasick (trie with transition links) can get matching substrings in linear time. Whoosh uses some fsa in their indexing as well but I did not investigate using their implementations (Elasticsearch uses a DAFSA under the hood I believe). I used the pyahocorasick library (<a href=\"https://pyahocorasick.readthedocs.io/en/latest/\" target=\"_blank\">https://pyahocorasick.readthedocs.io/en/latest/</a>) to build this fsa.</p>\n<ul>\n<li>Note that SpaCy noun chunks have root and extended forms, and do not consistently return the same chunks, but this is inherently resolved by the fsa and we can just extract the matching substrings</li>\n<li>To allow for exact phrase match (instead of substring within word match), pad each phrase with space at the start and end</li>\n</ul>\n<h6>Extract final phrases from parquet data and format into dataset</h6>\n<ul>\n<li>Run each description through the automaton and extract all matching phrases with their corresponding counts</li>\n<li>Counts are converted to ranks by scoring among all documents<ul>\n<li>Documents with same phrase counts get same rank but <em>lowered</em> e.g. for counts of 10, 5, 5, the ranks are 1, 3, 3 (not 1, 2, 2)</li></ul></li>\n<li>The phrase index is formatted as an integer list instead of a dictionary. Each stride of 3 represents a (phrase_id, rank, document_frequency)</li>\n<li>Polars hits array limit when saving the final output, I sort the phrase ranks to prioritize larger document frequencies and truncate the list length at 500 phrases per document</li>\n</ul>\n<h2>Weighted set cover algorithm</h2>\n<p>I approached the core problem as a variant of weighted set cover. 2 parts to the problem: </p>\n<ul>\n<li>calculating the weight of each phrase query (weight of a set is the sum of phrase query weights for each neighbour that contains the phrase)</li>\n<li>algorithm for picking sets</li>\n</ul>\n<p>Initially I spent little time on the weight formulation, using a harmonic series based ranking function, and spent the majority of time experimenting on approaches to set picking. Later on I kept to a simple greedy approach and focused on designing a simple, intuitive, and tunable weighting formula.</p>\n<h6>Phrase query weighting</h6>\n<p>I expressed the query weight as the interpolation of 2 weighting functions - rank based weighting and uniform.<br>\n<em>Rank based weighting</em>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F67bf8b4eb96e10a01738c82c4eebbe32%2FScreenshot%202024-07-25%20at%208.38.17PM.png?generation=1721911114726306&amp;alt=media\" alt=\"\"><br>\nDocuments with higher rank will have higher weight and the weight of all documents sums to 1.<br>\n<em>Uniform weighting</em>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F76e60cc9cf581680e714f87629eca6d3%2FScreenshot%202024-07-25%20at%208.39.40PM.png?generation=1721911199454127&amp;alt=media\" alt=\"\"><br>\nIntuitively this represents the scenario where for any search term, non-relevant documents are never in the index, so rank doesn’t matter since as long as the document has the term, it will always be relevant. As the index size gets smaller compared to the actual global size, this approximation gets better since it becomes increasingly unlikely that any non-relevant document is returned regardless of rank.<br>\n<em>Interpolated function</em>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F65cb63c5452e2bf298f3210ea05702ad%2FScreenshot%202024-07-25%20at%208.40.31PM.png?generation=1721911253003483&amp;alt=media\" alt=\"\"><br>\nThe alpha parameter represents the contribution of ranked and uniform components to the overall weight. Intuitively this represents the probability of false positives given an index size and global size. If the index and global are the same, then rank weights are directly equivalent to search ranks and the weight of a query term is directly equivalent to the ranks of the neighbour documents with that term. If the index is infinitesimally smaller than global, then rank doesn’t matter and the weight of a query term is only related to the number of neighbour documents with that term.</p>\n<h6>Set cover algorithm</h6>\n<p>Using weights from above, the simple greedy approach is to pick the term that provides the largest increase in total weight in every iteration. I introduce a new parameter, threshold, that represents the ‘saturation’ of weight for an item in the set. For example, if query term a contributes weight 1 to patent x, and threshold is 1, then any query term that also contributes to patent x doesn’t increase the total weight for x, since term a already returns x exactly.</p>\n<p>The algorithm loop is:</p>\n<ul>\n<li>Try each query term from all 50 neighbours and calculate the increase in weight for all neighbours, subject to threshold</li>\n<li>Pick the term with highest weight contribution, and remove it from candidates</li>\n<li>Repeat until maximum weight is reached (all neighbour weights reach threshold) or tokens exceed 50</li>\n</ul>\n<p>For faster computation I represent set cover as a bitset and only save the highest weight per unique bitset during weight calculation</p>\n<h2>Parameter optimization</h2>\n<p>In addition to alpha and saturation threshold, I have another threshold for filtering out low weightage terms in the initial calculation. These 3 parameters can be easily tuned, and what I tried with the limited time was Bayesian optimization, treating the full scoring against 1 validation index as 1 black box evaluation. In practice the effectiveness was severely limited due to bad validation indexes, query time limit making evaluations noisy, but I do believe that simple tuning is effective if these are resolved. The initial parameter bounds are:</p>\n<ul>\n<li>alpha (0, 0.8)</li>\n<li>weight threshold (0, 0.04)</li>\n<li>combination threshold (saturation) (0, 0.7)</li>\n</ul>\n<p>Here is a 3d visualization (credits: <a href=\"https://github.com/Yamanaka-Lab-TUAT/BOXVIA\" target=\"_blank\">https://github.com/Yamanaka-Lab-TUAT/BOXVIA</a>) of one BO set of trials against a validation index, with EI as acquisition function:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fb590afa1e271e15c3513e295e2f99e09%2FScreenshot%202024-07-25%20at%208.02.05PM.png?generation=1721911655634269&amp;alt=media\" alt=\"\"><br>\nNote that for this particular run it looks like fully uniform weighting for a range of combination thresholds is not far from the global maxima, and the weight threshold doesn't seem to matter much. </p>\n<h2>Closing notes</h2>\n<p>Tldr my solution consists of a heavy preprocessing pipeline to extract meaningful noun phrases, a greedy set cover algorithm with weighting that tries to approximate the sampling of the validation/test index from the global dataset, and some simple tuning. All the data and code are linked below and I look forward to any comments, feedback, suggestions, especially any mistakes or how I can improve. Thanks!</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/datasets/dilliontan/uspto-preprocessed/\" target=\"_blank\">https://www.kaggle.com/datasets/dilliontan/uspto-preprocessed/</a><br>\nSubmission notebook: <a href=\"https://www.kaggle.com/code/dilliontan/uspto-optimized-submission\" target=\"_blank\">https://www.kaggle.com/code/dilliontan/uspto-optimized-submission</a></p>\n<h6>Query examples</h6>\n<p>US-5516579-A<br>\n23 tokens covers all neighbours</p>\n<blockquote>\n  <p>detd:\"polymeric-phenolic-esters\" OR detd:\"predominant-dicarboxylic-acid\" OR detd:\"tmac-tma-bpda\" OR ti:\"solid-phase-polycondensation\" OR ab:\"rapid-molding-time\" OR clm:\"particulate-flame-retardants\" OR ab:\"especially-outstanding-toughness\" OR ab:\"ether-glycol-phthalate\" OR ab:\"new-economic-prepolymers\" OR ab:\"either-aliphatic-diacid\" OR clm:\"crystalline-polyester-imide\" OR ab:\"effective-reinforcing-amount\"</p>\n</blockquote>\n<p>US-8637833-B2<br>\n1 token covers all neighbours</p>\n<blockquote>\n  <p>detd:\"proton-delivery-efficiency\"</p>\n</blockquote>\n<p>US-11663219-B1<br>\n1 token covers all neighbours</p>\n<blockquote>\n  <p>detd:\"editable-buckets\"</p>\n</blockquote>\n<p>US-8047197-B1<br>\n49 tokens covers 29 neighbours</p>\n<blockquote>\n  <p>ab:\"tile-roofs\" OR clm:\"pointed-assembly\" OR ab:\"tile-roof\" OR ab:\"coupling-support-structure\" OR detd:\"screw-attached-mount\" OR clm:\"said-primary-barriers\" OR ab:\"phone-stand-assembly\" OR detd:\"vertically-orientated-bracket\" OR ti:\"food-cooling-assembly\" OR ti:\"roof-bracket-apparatus\" OR clm:\"fixedly-arranged-devices\" OR ti:\"rain-monitoring-system\" OR ab:\"tile-support-panels\" OR ab:\"consecutive-inclined-surfaces\" OR ti:\"laptop-support-assembly\" OR clm:\"bent-section-directing\" OR clm:\"convexly-arcuate-prominence\" OR ti:\"rod-holding-system\" OR clm:\"horizontally-orientated-members\" OR clm:\"telecommunications-rack-structure\" OR clm:\"middle-telescoping-member\" OR ti:\"tilt-table-system\" OR detd:\"first-elongate-spaces\" OR ti:\"door-painting-assembly\" OR ab:\"relatively-perpendicular-angle\"</p>\n</blockquote>\n<p>US-RE30747-E<br>\n49 tokens covers 42 neighbours</p>\n<blockquote>\n  <p>detd:\"muszik-et\" OR detd:\"structuring-substance\" OR detd:\"hectograph-blankets\" OR detd:\"fill-adhesive-composition\" OR detd:\"pregummed-wall-paper\" OR detd:\"outstanding-writability\" OR detd:\"decalcomania-coating\" OR detd:\"thermo-adhesive-tapes\" OR detd:\"di-dimethylcyclohexyl-adipate\" OR detd:\"single-plasticizers\" OR ti:\"polyvinyl-alcohol-deposits\" OR ti:\"decalcomania-adhesive\" OR ti:\"marine-dye-composition\" OR clm:\"water-dispersive-copolymer\" OR ti:\"pentavalent-vanadium-compound\" OR detd:\"example-starch-hydrate\" OR ab:\"medium-moisture-proof\" OR ti:\"coated-cork-composition\" OR detd:\"least-limitedsolubility\" OR detd:\"serial-tolerance\" OR ab:\"excellent-frangibility\" OR detd:\"difilcultto-remove\" OR detd:\"filly-name\" OR ab:\"stick-application\" OR clm:\"octyl-terephthalic-acid\"</p>\n</blockquote>\n<p>US-10013212-B2<br>\n49 tokens covers 38 neighbours</p>\n<blockquote>\n  <p>detd:\"intrusion-detection-database\" OR detd:\"avalon-memory-mapped\" OR clm:\"programmable-hardware-units\" OR ab:\"block-rams\" OR detd:\"reconfigurable-gate-arrays\" OR detd:\"stabilizing-logic-devices\" OR clm:\"programmable-logic-processor\" OR detd:\"partially-reconfigurable-regions\" OR ti:\"heterogeneous-memory-primitives\" OR ab:\"guardband-voltages\" OR ti:\"gpu-virtualisation\" OR ab:\"least-arbitration-rule\" OR ab:\"configuration-programming-block\" OR clm:\"variable-size-ratio\" OR ti:\"immediate-value-network\" OR ab:\"iii-cache-data\" OR ti:\"fast-dynamic-change\" OR ti:\"matrix-storage-method\" OR clm:\"asynchronous-handshaking-scheme\" OR ti:\"hierarchical-sort-acceleration\" OR clm:\"execution-complete-flag\" OR clm:\"coherent-interconnection-protocol\" OR clm:\"intermediate-module-channel\" OR ab:\"correct-storage-address\" OR detd:\"problematic-memory-blocks\"</p>\n</blockquote>",
  "messages": [
    {
      "id": 2935633,
      "postDate": "2024-07-25T13:01:24.983Z",
      "content": "<p>My initial solution was unigram combinations of all fields, but after discovering the ‘magic character’* in phrase search and getting good initial results, I pivoted to using bigrams and trigrams only. Ultimately it wasn’t a good decision for the scoring and my only consolation is that the majority of bigrams/trigrams query terms in my final solution are quite meaningful. It’s also quite cool when target patents can be matched by a single bigram/trigram.</p>\n<p>I’ll go into detail about 3 main parts of my final solution:</p>\n<p><a href=\"#preprocessing\">Preprocessing</a><br>\n<a href=\"#weighted-set-cover-algorithm\">Weighted set cover algorithm</a><br>\n<a href=\"#parameter-optimization\">Parameter optimization</a></p>\n<p>*Side note:<br>\nThe ‘magic character’ that I’m referring to is the usage of special characters to join words in phrases, which is the only trick I used in my solution. The character I used to join phrases was ‘-’. The token savings of this trick in my solution was actually not that significant in practice, due to the following reasons:</p>\n<ul>\n<li>Because I want to use meaningful bigrams and trigrams in all fields, it is more important to find ngrams that can target span patent groups. In a significant proportion of patent groups I think there actually aren’t such bigrams/trigrams, so it imposes an upper bound on the possible cover with 50 tokens. In some patent groups however it does work very nicely and you can actually find a bigram/trigram that can cover all 50 neighbors without false positives (in such cases the trick also doesn’t matter since we only need a few tokens)</li>\n<li>Phrase search in description has VERY expensive query time and a major issue I faced in the last 2 weeks after I created my final phrase dataset. I came to realise that some bigrams/trigrams can even take more than 20s on their own and I had to implement multiple changes to get my notebook to score reliably. These findings are documented in a separate post <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522303\" target=\"_blank\">here</a></li>\n<li>Phrase search in Whoosh had some inconsistencies such as the one described <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/519755#2919472\" target=\"_blank\">here</a>, that is complex to account for in the preprocessing pipeline</li>\n</ul>\n<p>In summary, phrase search in title, abstract and claims is quite stable, but in description field … <em>here be dragons</em>. At the same time description field accounts for majority of the query results for this approach and having only title, claims and abstract would score very poorly.</p>\n<h2>Preprocessing</h2>\n<p>The goal of the preprocessing pipeline is to generate phrase ranks for all patents such that for each target patent, we can fetch ranked phrases for each neighbour and assemble them together quickly. A finite state automaton built using Aho-Corasick is the key part of this pipeline.</p>\n<h6>SpaCy NLP + CountVectorizer</h6>\n<p>To extract noun chunks I used SpaCy’s nlp pipeline inside a batched CountVectorizer with optimizations like</p>\n<ul>\n<li>disabling lemmatizer and ner in nlp pipeline</li>\n<li>analyze only first 1 million characters in description text (SpaCy’s default maximum length for nlp pipeline is 1 million)</li>\n<li>binary True and max_df of 0.1 for CountVectorizer</li>\n<li>truncating or dropping noun chunks with special characters</li>\n<li>dropping noun chunks with count &gt; 150 after combining batches</li>\n<li>clean the combined vocabulary, check for length, Shannon entropy, etc.  </li>\n</ul>\n<p>Final output is a vocabulary of phrases (the count obtained here is an estimate, the true count will be obtained later through the fsa)</p>\n<h6>Build fsa from vocabulary</h6>\n<p>A finite state automaton using Aho-Corasick (trie with transition links) can get matching substrings in linear time. Whoosh uses some fsa in their indexing as well but I did not investigate using their implementations (Elasticsearch uses a DAFSA under the hood I believe). I used the pyahocorasick library (<a href=\"https://pyahocorasick.readthedocs.io/en/latest/\" target=\"_blank\">https://pyahocorasick.readthedocs.io/en/latest/</a>) to build this fsa.</p>\n<ul>\n<li>Note that SpaCy noun chunks have root and extended forms, and do not consistently return the same chunks, but this is inherently resolved by the fsa and we can just extract the matching substrings</li>\n<li>To allow for exact phrase match (instead of substring within word match), pad each phrase with space at the start and end</li>\n</ul>\n<h6>Extract final phrases from parquet data and format into dataset</h6>\n<ul>\n<li>Run each description through the automaton and extract all matching phrases with their corresponding counts</li>\n<li>Counts are converted to ranks by scoring among all documents<ul>\n<li>Documents with same phrase counts get same rank but <em>lowered</em> e.g. for counts of 10, 5, 5, the ranks are 1, 3, 3 (not 1, 2, 2)</li></ul></li>\n<li>The phrase index is formatted as an integer list instead of a dictionary. Each stride of 3 represents a (phrase_id, rank, document_frequency)</li>\n<li>Polars hits array limit when saving the final output, I sort the phrase ranks to prioritize larger document frequencies and truncate the list length at 500 phrases per document</li>\n</ul>\n<h2>Weighted set cover algorithm</h2>\n<p>I approached the core problem as a variant of weighted set cover. 2 parts to the problem: </p>\n<ul>\n<li>calculating the weight of each phrase query (weight of a set is the sum of phrase query weights for each neighbour that contains the phrase)</li>\n<li>algorithm for picking sets</li>\n</ul>\n<p>Initially I spent little time on the weight formulation, using a harmonic series based ranking function, and spent the majority of time experimenting on approaches to set picking. Later on I kept to a simple greedy approach and focused on designing a simple, intuitive, and tunable weighting formula.</p>\n<h6>Phrase query weighting</h6>\n<p>I expressed the query weight as the interpolation of 2 weighting functions - rank based weighting and uniform.<br>\n<em>Rank based weighting</em>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F67bf8b4eb96e10a01738c82c4eebbe32%2FScreenshot%202024-07-25%20at%208.38.17PM.png?generation=1721911114726306&amp;alt=media\" alt=\"\"><br>\nDocuments with higher rank will have higher weight and the weight of all documents sums to 1.<br>\n<em>Uniform weighting</em>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F76e60cc9cf581680e714f87629eca6d3%2FScreenshot%202024-07-25%20at%208.39.40PM.png?generation=1721911199454127&amp;alt=media\" alt=\"\"><br>\nIntuitively this represents the scenario where for any search term, non-relevant documents are never in the index, so rank doesn’t matter since as long as the document has the term, it will always be relevant. As the index size gets smaller compared to the actual global size, this approximation gets better since it becomes increasingly unlikely that any non-relevant document is returned regardless of rank.<br>\n<em>Interpolated function</em>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F65cb63c5452e2bf298f3210ea05702ad%2FScreenshot%202024-07-25%20at%208.40.31PM.png?generation=1721911253003483&amp;alt=media\" alt=\"\"><br>\nThe alpha parameter represents the contribution of ranked and uniform components to the overall weight. Intuitively this represents the probability of false positives given an index size and global size. If the index and global are the same, then rank weights are directly equivalent to search ranks and the weight of a query term is directly equivalent to the ranks of the neighbour documents with that term. If the index is infinitesimally smaller than global, then rank doesn’t matter and the weight of a query term is only related to the number of neighbour documents with that term.</p>\n<h6>Set cover algorithm</h6>\n<p>Using weights from above, the simple greedy approach is to pick the term that provides the largest increase in total weight in every iteration. I introduce a new parameter, threshold, that represents the ‘saturation’ of weight for an item in the set. For example, if query term a contributes weight 1 to patent x, and threshold is 1, then any query term that also contributes to patent x doesn’t increase the total weight for x, since term a already returns x exactly.</p>\n<p>The algorithm loop is:</p>\n<ul>\n<li>Try each query term from all 50 neighbours and calculate the increase in weight for all neighbours, subject to threshold</li>\n<li>Pick the term with highest weight contribution, and remove it from candidates</li>\n<li>Repeat until maximum weight is reached (all neighbour weights reach threshold) or tokens exceed 50</li>\n</ul>\n<p>For faster computation I represent set cover as a bitset and only save the highest weight per unique bitset during weight calculation</p>\n<h2>Parameter optimization</h2>\n<p>In addition to alpha and saturation threshold, I have another threshold for filtering out low weightage terms in the initial calculation. These 3 parameters can be easily tuned, and what I tried with the limited time was Bayesian optimization, treating the full scoring against 1 validation index as 1 black box evaluation. In practice the effectiveness was severely limited due to bad validation indexes, query time limit making evaluations noisy, but I do believe that simple tuning is effective if these are resolved. The initial parameter bounds are:</p>\n<ul>\n<li>alpha (0, 0.8)</li>\n<li>weight threshold (0, 0.04)</li>\n<li>combination threshold (saturation) (0, 0.7)</li>\n</ul>\n<p>Here is a 3d visualization (credits: <a href=\"https://github.com/Yamanaka-Lab-TUAT/BOXVIA\" target=\"_blank\">https://github.com/Yamanaka-Lab-TUAT/BOXVIA</a>) of one BO set of trials against a validation index, with EI as acquisition function:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fb590afa1e271e15c3513e295e2f99e09%2FScreenshot%202024-07-25%20at%208.02.05PM.png?generation=1721911655634269&amp;alt=media\" alt=\"\"><br>\nNote that for this particular run it looks like fully uniform weighting for a range of combination thresholds is not far from the global maxima, and the weight threshold doesn't seem to matter much. </p>\n<h2>Closing notes</h2>\n<p>Tldr my solution consists of a heavy preprocessing pipeline to extract meaningful noun phrases, a greedy set cover algorithm with weighting that tries to approximate the sampling of the validation/test index from the global dataset, and some simple tuning. All the data and code are linked below and I look forward to any comments, feedback, suggestions, especially any mistakes or how I can improve. Thanks!</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/datasets/dilliontan/uspto-preprocessed/\" target=\"_blank\">https://www.kaggle.com/datasets/dilliontan/uspto-preprocessed/</a><br>\nSubmission notebook: <a href=\"https://www.kaggle.com/code/dilliontan/uspto-optimized-submission\" target=\"_blank\">https://www.kaggle.com/code/dilliontan/uspto-optimized-submission</a></p>\n<h6>Query examples</h6>\n<p>US-5516579-A<br>\n23 tokens covers all neighbours</p>\n<blockquote>\n  <p>detd:\"polymeric-phenolic-esters\" OR detd:\"predominant-dicarboxylic-acid\" OR detd:\"tmac-tma-bpda\" OR ti:\"solid-phase-polycondensation\" OR ab:\"rapid-molding-time\" OR clm:\"particulate-flame-retardants\" OR ab:\"especially-outstanding-toughness\" OR ab:\"ether-glycol-phthalate\" OR ab:\"new-economic-prepolymers\" OR ab:\"either-aliphatic-diacid\" OR clm:\"crystalline-polyester-imide\" OR ab:\"effective-reinforcing-amount\"</p>\n</blockquote>\n<p>US-8637833-B2<br>\n1 token covers all neighbours</p>\n<blockquote>\n  <p>detd:\"proton-delivery-efficiency\"</p>\n</blockquote>\n<p>US-11663219-B1<br>\n1 token covers all neighbours</p>\n<blockquote>\n  <p>detd:\"editable-buckets\"</p>\n</blockquote>\n<p>US-8047197-B1<br>\n49 tokens covers 29 neighbours</p>\n<blockquote>\n  <p>ab:\"tile-roofs\" OR clm:\"pointed-assembly\" OR ab:\"tile-roof\" OR ab:\"coupling-support-structure\" OR detd:\"screw-attached-mount\" OR clm:\"said-primary-barriers\" OR ab:\"phone-stand-assembly\" OR detd:\"vertically-orientated-bracket\" OR ti:\"food-cooling-assembly\" OR ti:\"roof-bracket-apparatus\" OR clm:\"fixedly-arranged-devices\" OR ti:\"rain-monitoring-system\" OR ab:\"tile-support-panels\" OR ab:\"consecutive-inclined-surfaces\" OR ti:\"laptop-support-assembly\" OR clm:\"bent-section-directing\" OR clm:\"convexly-arcuate-prominence\" OR ti:\"rod-holding-system\" OR clm:\"horizontally-orientated-members\" OR clm:\"telecommunications-rack-structure\" OR clm:\"middle-telescoping-member\" OR ti:\"tilt-table-system\" OR detd:\"first-elongate-spaces\" OR ti:\"door-painting-assembly\" OR ab:\"relatively-perpendicular-angle\"</p>\n</blockquote>\n<p>US-RE30747-E<br>\n49 tokens covers 42 neighbours</p>\n<blockquote>\n  <p>detd:\"muszik-et\" OR detd:\"structuring-substance\" OR detd:\"hectograph-blankets\" OR detd:\"fill-adhesive-composition\" OR detd:\"pregummed-wall-paper\" OR detd:\"outstanding-writability\" OR detd:\"decalcomania-coating\" OR detd:\"thermo-adhesive-tapes\" OR detd:\"di-dimethylcyclohexyl-adipate\" OR detd:\"single-plasticizers\" OR ti:\"polyvinyl-alcohol-deposits\" OR ti:\"decalcomania-adhesive\" OR ti:\"marine-dye-composition\" OR clm:\"water-dispersive-copolymer\" OR ti:\"pentavalent-vanadium-compound\" OR detd:\"example-starch-hydrate\" OR ab:\"medium-moisture-proof\" OR ti:\"coated-cork-composition\" OR detd:\"least-limitedsolubility\" OR detd:\"serial-tolerance\" OR ab:\"excellent-frangibility\" OR detd:\"difilcultto-remove\" OR detd:\"filly-name\" OR ab:\"stick-application\" OR clm:\"octyl-terephthalic-acid\"</p>\n</blockquote>\n<p>US-10013212-B2<br>\n49 tokens covers 38 neighbours</p>\n<blockquote>\n  <p>detd:\"intrusion-detection-database\" OR detd:\"avalon-memory-mapped\" OR clm:\"programmable-hardware-units\" OR ab:\"block-rams\" OR detd:\"reconfigurable-gate-arrays\" OR detd:\"stabilizing-logic-devices\" OR clm:\"programmable-logic-processor\" OR detd:\"partially-reconfigurable-regions\" OR ti:\"heterogeneous-memory-primitives\" OR ab:\"guardband-voltages\" OR ti:\"gpu-virtualisation\" OR ab:\"least-arbitration-rule\" OR ab:\"configuration-programming-block\" OR clm:\"variable-size-ratio\" OR ti:\"immediate-value-network\" OR ab:\"iii-cache-data\" OR ti:\"fast-dynamic-change\" OR ti:\"matrix-storage-method\" OR clm:\"asynchronous-handshaking-scheme\" OR ti:\"hierarchical-sort-acceleration\" OR clm:\"execution-complete-flag\" OR clm:\"coherent-interconnection-protocol\" OR clm:\"intermediate-module-channel\" OR ab:\"correct-storage-address\" OR detd:\"problematic-memory-blocks\"</p>\n</blockquote>",
      "rawMarkdown": "My initial solution was unigram combinations of all fields, but after discovering the ‘magic character’* in phrase search and getting good initial results, I pivoted to using bigrams and trigrams only. Ultimately it wasn’t a good decision for the scoring and my only consolation is that the majority of bigrams/trigrams query terms in my final solution are quite meaningful. It’s also quite cool when target patents can be matched by a single bigram/trigram.\n\nI’ll go into detail about 3 main parts of my final solution:\n\n[Preprocessing](#preprocessing)\n[Weighted set cover algorithm](#weighted-set-cover-algorithm)\n[Parameter optimization](#parameter-optimization)\n\n*Side note:\nThe ‘magic character’ that I’m referring to is the usage of special characters to join words in phrases, which is the only trick I used in my solution. The character I used to join phrases was ‘-’. The token savings of this trick in my solution was actually not that significant in practice, due to the following reasons:\n- Because I want to use meaningful bigrams and trigrams in all fields, it is more important to find ngrams that can target span patent groups. In a significant proportion of patent groups I think there actually aren’t such bigrams/trigrams, so it imposes an upper bound on the possible cover with 50 tokens. In some patent groups however it does work very nicely and you can actually find a bigram/trigram that can cover all 50 neighbors without false positives (in such cases the trick also doesn’t matter since we only need a few tokens)\n- Phrase search in description has VERY expensive query time and a major issue I faced in the last 2 weeks after I created my final phrase dataset. I came to realise that some bigrams/trigrams can even take more than 20s on their own and I had to implement multiple changes to get my notebook to score reliably. These findings are documented in a separate post [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522303)\n- Phrase search in Whoosh had some inconsistencies such as the one described [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/519755#2919472), that is complex to account for in the preprocessing pipeline\n\nIn summary, phrase search in title, abstract and claims is quite stable, but in description field … *here be dragons*. At the same time description field accounts for majority of the query results for this approach and having only title, claims and abstract would score very poorly.\n\n## Preprocessing\nThe goal of the preprocessing pipeline is to generate phrase ranks for all patents such that for each target patent, we can fetch ranked phrases for each neighbour and assemble them together quickly. A finite state automaton built using Aho-Corasick is the key part of this pipeline.\n\n###### SpaCy NLP + CountVectorizer\nTo extract noun chunks I used SpaCy’s nlp pipeline inside a batched CountVectorizer with optimizations like\n- disabling lemmatizer and ner in nlp pipeline\n- analyze only first 1 million characters in description text (SpaCy’s default maximum length for nlp pipeline is 1 million)\n- binary True and max_df of 0.1 for CountVectorizer\n- truncating or dropping noun chunks with special characters\n- dropping noun chunks with count > 150 after combining batches\n- clean the combined vocabulary, check for length, Shannon entropy, etc.  \n\nFinal output is a vocabulary of phrases (the count obtained here is an estimate, the true count will be obtained later through the fsa)\n\n###### Build fsa from vocabulary\nA finite state automaton using Aho-Corasick (trie with transition links) can get matching substrings in linear time. Whoosh uses some fsa in their indexing as well but I did not investigate using their implementations (Elasticsearch uses a DAFSA under the hood I believe). I used the pyahocorasick library (https://pyahocorasick.readthedocs.io/en/latest/) to build this fsa.\n- Note that SpaCy noun chunks have root and extended forms, and do not consistently return the same chunks, but this is inherently resolved by the fsa and we can just extract the matching substrings\n- To allow for exact phrase match (instead of substring within word match), pad each phrase with space at the start and end\n\n###### Extract final phrases from parquet data and format into dataset\n- Run each description through the automaton and extract all matching phrases with their corresponding counts\n- Counts are converted to ranks by scoring among all documents\n  - Documents with same phrase counts get same rank but *lowered* e.g. for counts of 10, 5, 5, the ranks are 1, 3, 3 (not 1, 2, 2)\n- The phrase index is formatted as an integer list instead of a dictionary. Each stride of 3 represents a (phrase_id, rank, document_frequency)\n- Polars hits array limit when saving the final output, I sort the phrase ranks to prioritize larger document frequencies and truncate the list length at 500 phrases per document\n\n## Weighted set cover algorithm\nI approached the core problem as a variant of weighted set cover. 2 parts to the problem: \n- calculating the weight of each phrase query (weight of a set is the sum of phrase query weights for each neighbour that contains the phrase)\n- algorithm for picking sets\n\nInitially I spent little time on the weight formulation, using a harmonic series based ranking function, and spent the majority of time experimenting on approaches to set picking. Later on I kept to a simple greedy approach and focused on designing a simple, intuitive, and tunable weighting formula.\n\n###### Phrase query weighting\nI expressed the query weight as the interpolation of 2 weighting functions - rank based weighting and uniform.\n*Rank based weighting*:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F67bf8b4eb96e10a01738c82c4eebbe32%2FScreenshot%202024-07-25%20at%208.38.17PM.png?generation=1721911114726306&alt=media)\nDocuments with higher rank will have higher weight and the weight of all documents sums to 1.\n*Uniform weighting*:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F76e60cc9cf581680e714f87629eca6d3%2FScreenshot%202024-07-25%20at%208.39.40PM.png?generation=1721911199454127&alt=media)\nIntuitively this represents the scenario where for any search term, non-relevant documents are never in the index, so rank doesn’t matter since as long as the document has the term, it will always be relevant. As the index size gets smaller compared to the actual global size, this approximation gets better since it becomes increasingly unlikely that any non-relevant document is returned regardless of rank.\n*Interpolated function*:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F65cb63c5452e2bf298f3210ea05702ad%2FScreenshot%202024-07-25%20at%208.40.31PM.png?generation=1721911253003483&alt=media)\nThe alpha parameter represents the contribution of ranked and uniform components to the overall weight. Intuitively this represents the probability of false positives given an index size and global size. If the index and global are the same, then rank weights are directly equivalent to search ranks and the weight of a query term is directly equivalent to the ranks of the neighbour documents with that term. If the index is infinitesimally smaller than global, then rank doesn’t matter and the weight of a query term is only related to the number of neighbour documents with that term.\n\n###### Set cover algorithm\nUsing weights from above, the simple greedy approach is to pick the term that provides the largest increase in total weight in every iteration. I introduce a new parameter, threshold, that represents the ‘saturation’ of weight for an item in the set. For example, if query term a contributes weight 1 to patent x, and threshold is 1, then any query term that also contributes to patent x doesn’t increase the total weight for x, since term a already returns x exactly.\n\nThe algorithm loop is:\n- Try each query term from all 50 neighbours and calculate the increase in weight for all neighbours, subject to threshold\n- Pick the term with highest weight contribution, and remove it from candidates\n- Repeat until maximum weight is reached (all neighbour weights reach threshold) or tokens exceed 50\n\nFor faster computation I represent set cover as a bitset and only save the highest weight per unique bitset during weight calculation\n\n## Parameter optimization\nIn addition to alpha and saturation threshold, I have another threshold for filtering out low weightage terms in the initial calculation. These 3 parameters can be easily tuned, and what I tried with the limited time was Bayesian optimization, treating the full scoring against 1 validation index as 1 black box evaluation. In practice the effectiveness was severely limited due to bad validation indexes, query time limit making evaluations noisy, but I do believe that simple tuning is effective if these are resolved. The initial parameter bounds are:\n- alpha (0, 0.8)\n- weight threshold (0, 0.04)\n- combination threshold (saturation) (0, 0.7)\n\nHere is a 3d visualization (credits: https://github.com/Yamanaka-Lab-TUAT/BOXVIA) of one BO set of trials against a validation index, with EI as acquisition function:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fb590afa1e271e15c3513e295e2f99e09%2FScreenshot%202024-07-25%20at%208.02.05PM.png?generation=1721911655634269&alt=media)\nNote that for this particular run it looks like fully uniform weighting for a range of combination thresholds is not far from the global maxima, and the weight threshold doesn't seem to matter much. \n\n## Closing notes\nTldr my solution consists of a heavy preprocessing pipeline to extract meaningful noun phrases, a greedy set cover algorithm with weighting that tries to approximate the sampling of the validation/test index from the global dataset, and some simple tuning. All the data and code are linked below and I look forward to any comments, feedback, suggestions, especially any mistakes or how I can improve. Thanks!\n\n\nDataset: https://www.kaggle.com/datasets/dilliontan/uspto-preprocessed/\nSubmission notebook: https://www.kaggle.com/code/dilliontan/uspto-optimized-submission\n\n###### Query examples\n\nUS-5516579-A\n23 tokens covers all neighbours\n>detd:\"polymeric-phenolic-esters\" OR detd:\"predominant-dicarboxylic-acid\" OR detd:\"tmac-tma-bpda\" OR ti:\"solid-phase-polycondensation\" OR ab:\"rapid-molding-time\" OR clm:\"particulate-flame-retardants\" OR ab:\"especially-outstanding-toughness\" OR ab:\"ether-glycol-phthalate\" OR ab:\"new-economic-prepolymers\" OR ab:\"either-aliphatic-diacid\" OR clm:\"crystalline-polyester-imide\" OR ab:\"effective-reinforcing-amount\"\n\nUS-8637833-B2\n1 token covers all neighbours\n>detd:\"proton-delivery-efficiency\"\n\nUS-11663219-B1\n1 token covers all neighbours\n>detd:\"editable-buckets\"\n\nUS-8047197-B1\n49 tokens covers 29 neighbours\n>ab:\"tile-roofs\" OR clm:\"pointed-assembly\" OR ab:\"tile-roof\" OR ab:\"coupling-support-structure\" OR detd:\"screw-attached-mount\" OR clm:\"said-primary-barriers\" OR ab:\"phone-stand-assembly\" OR detd:\"vertically-orientated-bracket\" OR ti:\"food-cooling-assembly\" OR ti:\"roof-bracket-apparatus\" OR clm:\"fixedly-arranged-devices\" OR ti:\"rain-monitoring-system\" OR ab:\"tile-support-panels\" OR ab:\"consecutive-inclined-surfaces\" OR ti:\"laptop-support-assembly\" OR clm:\"bent-section-directing\" OR clm:\"convexly-arcuate-prominence\" OR ti:\"rod-holding-system\" OR clm:\"horizontally-orientated-members\" OR clm:\"telecommunications-rack-structure\" OR clm:\"middle-telescoping-member\" OR ti:\"tilt-table-system\" OR detd:\"first-elongate-spaces\" OR ti:\"door-painting-assembly\" OR ab:\"relatively-perpendicular-angle\"\n\nUS-RE30747-E\n49 tokens covers 42 neighbours\n>detd:\"muszik-et\" OR detd:\"structuring-substance\" OR detd:\"hectograph-blankets\" OR detd:\"fill-adhesive-composition\" OR detd:\"pregummed-wall-paper\" OR detd:\"outstanding-writability\" OR detd:\"decalcomania-coating\" OR detd:\"thermo-adhesive-tapes\" OR detd:\"di-dimethylcyclohexyl-adipate\" OR detd:\"single-plasticizers\" OR ti:\"polyvinyl-alcohol-deposits\" OR ti:\"decalcomania-adhesive\" OR ti:\"marine-dye-composition\" OR clm:\"water-dispersive-copolymer\" OR ti:\"pentavalent-vanadium-compound\" OR detd:\"example-starch-hydrate\" OR ab:\"medium-moisture-proof\" OR ti:\"coated-cork-composition\" OR detd:\"least-limitedsolubility\" OR detd:\"serial-tolerance\" OR ab:\"excellent-frangibility\" OR detd:\"difilcultto-remove\" OR detd:\"filly-name\" OR ab:\"stick-application\" OR clm:\"octyl-terephthalic-acid\"\n\nUS-10013212-B2\n49 tokens covers 38 neighbours\n>detd:\"intrusion-detection-database\" OR detd:\"avalon-memory-mapped\" OR clm:\"programmable-hardware-units\" OR ab:\"block-rams\" OR detd:\"reconfigurable-gate-arrays\" OR detd:\"stabilizing-logic-devices\" OR clm:\"programmable-logic-processor\" OR detd:\"partially-reconfigurable-regions\" OR ti:\"heterogeneous-memory-primitives\" OR ab:\"guardband-voltages\" OR ti:\"gpu-virtualisation\" OR ab:\"least-arbitration-rule\" OR ab:\"configuration-programming-block\" OR clm:\"variable-size-ratio\" OR ti:\"immediate-value-network\" OR ab:\"iii-cache-data\" OR ti:\"fast-dynamic-change\" OR ti:\"matrix-storage-method\" OR clm:\"asynchronous-handshaking-scheme\" OR ti:\"hierarchical-sort-acceleration\" OR clm:\"execution-complete-flag\" OR clm:\"coherent-interconnection-protocol\" OR clm:\"intermediate-module-channel\" OR ab:\"correct-storage-address\" OR detd:\"problematic-memory-blocks\"",
      "votes": 17
    },
    {
      "id": 2940830,
      "postDate": "2024-07-30T13:47:52.447Z",
      "content": "<p>Congratulations on earning a gold medal and achieving 12th place in the competition! <a href=\"https://www.kaggle.com/dilliontan\" target=\"_blank\">@dilliontan</a> <br>\nYour approach to using bigrams and trigrams, along with the detailed explanation of your preprocessing pipeline and weighted set cover algorithm, is impressive. It’s fascinating to see how you tackled the challenge with a combination of meaningful n-grams and advanced techniques like Aho-Corasick for efficient substring matching. Your insights into the nuances of phrase search, particularly in the description field, and the impact of special characters on token savings add valuable depth to your solution. Thanks for sharing your detailed methodology and the links to your dataset and code—these will undoubtedly benefit others looking to refine their own approaches in similar tasks.</p>",
      "rawMarkdown": "Congratulations on earning a gold medal and achieving 12th place in the competition! @dilliontan \nYour approach to using bigrams and trigrams, along with the detailed explanation of your preprocessing pipeline and weighted set cover algorithm, is impressive. It’s fascinating to see how you tackled the challenge with a combination of meaningful n-grams and advanced techniques like Aho-Corasick for efficient substring matching. Your insights into the nuances of phrase search, particularly in the description field, and the impact of special characters on token savings add valuable depth to your solution. Thanks for sharing your detailed methodology and the links to your dataset and code—these will undoubtedly benefit others looking to refine their own approaches in similar tasks."
    },
    {
      "id": 2935973,
      "postDate": "2024-07-25T16:53:50.033Z",
      "content": "<p>Thanks for sharing! Would you mind tagging this as a solution writeup so other people can access it from the leaderboard as well?</p>",
      "rawMarkdown": "Thanks for sharing! Would you mind tagging this as a solution writeup so other people can access it from the leaderboard as well?",
      "replies": [
        {
          "id": 2936461,
          "postDate": "2024-07-26T06:59:28.287Z",
          "content": "<p>for sure, just updated</p>\n<p>thanks for the reminder, I'm still new at navigating Kaggle!</p>",
          "rawMarkdown": "for sure, just updated\n\nthanks for the reminder, I'm still new at navigating Kaggle!"
        }
      ]
    },
    {
      "id": 2936670,
      "postDate": "2024-07-26T11:25:39.610Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2940830,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T13:47:52.447000",
      "content": "<p>Congratulations on earning a gold medal and achieving 12th place in the competition! <a href=\"https://www.kaggle.com/dilliontan\" target=\"_blank\">@dilliontan</a> <br>\nYour approach to using bigrams and trigrams, along with the detailed explanation of your preprocessing pipeline and weighted set cover algorithm, is impressive. It’s fascinating to see how you tackled the challenge with a combination of meaningful n-grams and advanced techniques like Aho-Corasick for efficient substring matching. Your insights into the nuances of phrase search, particularly in the description field, and the impact of special characters on token savings add valuable depth to your solution. Thanks for sharing your detailed methodology and the links to your dataset and code—these will undoubtedly benefit others looking to refine their own approaches in similar tasks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935973,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2024-07-25T16:53:50.033000",
      "content": "<p>Thanks for sharing! Would you mind tagging this as a solution writeup so other people can access it from the leaderboard as well?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2936461,
          "author_name": "dt",
          "author_url": "",
          "post_date": "2024-07-26T06:59:28.287000",
          "content": "<p>for sure, just updated</p>\n<p>thanks for the reminder, I'm still new at navigating Kaggle!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2936670,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-26T11:25:39.610000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935633": "My initial solution was unigram combinations of all fields, but after discovering the ‘magic character’* in phrase search and getting good initial results, I pivoted to using bigrams and trigrams only. Ultimately it wasn’t a good decision for the scoring and my only consolation is that the majority of bigrams/trigrams query terms in my final solution are quite meaningful. It’s also quite cool when target patents can be matched by a single bigram/trigram.\n\nI’ll go into detail about 3 main parts of my final solution:\n\n[Preprocessing](#preprocessing)\n[Weighted set cover algorithm](#weighted-set-cover-algorithm)\n[Parameter optimization](#parameter-optimization)\n\n*Side note:\nThe ‘magic character’ that I’m referring to is the usage of special characters to join words in phrases, which is the only trick I used in my solution. The character I used to join phrases was ‘-’. The token savings of this trick in my solution was actually not that significant in practice, due to the following reasons:\n- Because I want to use meaningful bigrams and trigrams in all fields, it is more important to find ngrams that can target span patent groups. In a significant proportion of patent groups I think there actually aren’t such bigrams/trigrams, so it imposes an upper bound on the possible cover with 50 tokens. In some patent groups however it does work very nicely and you can actually find a bigram/trigram that can cover all 50 neighbors without false positives (in such cases the trick also doesn’t matter since we only need a few tokens)\n- Phrase search in description has VERY expensive query time and a major issue I faced in the last 2 weeks after I created my final phrase dataset. I came to realise that some bigrams/trigrams can even take more than 20s on their own and I had to implement multiple changes to get my notebook to score reliably. These findings are documented in a separate post [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522303)\n- Phrase search in Whoosh had some inconsistencies such as the one described [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/519755#2919472), that is complex to account for in the preprocessing pipeline\n\nIn summary, phrase search in title, abstract and claims is quite stable, but in description field … *here be dragons*. At the same time description field accounts for majority of the query results for this approach and having only title, claims and abstract would score very poorly.\n\n## Preprocessing\nThe goal of the preprocessing pipeline is to generate phrase ranks for all patents such that for each target patent, we can fetch ranked phrases for each neighbour and assemble them together quickly. A finite state automaton built using Aho-Corasick is the key part of this pipeline.\n\n###### SpaCy NLP + CountVectorizer\nTo extract noun chunks I used SpaCy’s nlp pipeline inside a batched CountVectorizer with optimizations like\n- disabling lemmatizer and ner in nlp pipeline\n- analyze only first 1 million characters in description text (SpaCy’s default maximum length for nlp pipeline is 1 million)\n- binary True and max_df of 0.1 for CountVectorizer\n- truncating or dropping noun chunks with special characters\n- dropping noun chunks with count > 150 after combining batches\n- clean the combined vocabulary, check for length, Shannon entropy, etc.  \n\nFinal output is a vocabulary of phrases (the count obtained here is an estimate, the true count will be obtained later through the fsa)\n\n###### Build fsa from vocabulary\nA finite state automaton using Aho-Corasick (trie with transition links) can get matching substrings in linear time. Whoosh uses some fsa in their indexing as well but I did not investigate using their implementations (Elasticsearch uses a DAFSA under the hood I believe). I used the pyahocorasick library (https://pyahocorasick.readthedocs.io/en/latest/) to build this fsa.\n- Note that SpaCy noun chunks have root and extended forms, and do not consistently return the same chunks, but this is inherently resolved by the fsa and we can just extract the matching substrings\n- To allow for exact phrase match (instead of substring within word match), pad each phrase with space at the start and end\n\n###### Extract final phrases from parquet data and format into dataset\n- Run each description through the automaton and extract all matching phrases with their corresponding counts\n- Counts are converted to ranks by scoring among all documents\n  - Documents with same phrase counts get same rank but *lowered* e.g. for counts of 10, 5, 5, the ranks are 1, 3, 3 (not 1, 2, 2)\n- The phrase index is formatted as an integer list instead of a dictionary. Each stride of 3 represents a (phrase_id, rank, document_frequency)\n- Polars hits array limit when saving the final output, I sort the phrase ranks to prioritize larger document frequencies and truncate the list length at 500 phrases per document\n\n## Weighted set cover algorithm\nI approached the core problem as a variant of weighted set cover. 2 parts to the problem: \n- calculating the weight of each phrase query (weight of a set is the sum of phrase query weights for each neighbour that contains the phrase)\n- algorithm for picking sets\n\nInitially I spent little time on the weight formulation, using a harmonic series based ranking function, and spent the majority of time experimenting on approaches to set picking. Later on I kept to a simple greedy approach and focused on designing a simple, intuitive, and tunable weighting formula.\n\n###### Phrase query weighting\nI expressed the query weight as the interpolation of 2 weighting functions - rank based weighting and uniform.\n*Rank based weighting*:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F67bf8b4eb96e10a01738c82c4eebbe32%2FScreenshot%202024-07-25%20at%208.38.17PM.png?generation=1721911114726306&alt=media)\nDocuments with higher rank will have higher weight and the weight of all documents sums to 1.\n*Uniform weighting*:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F76e60cc9cf581680e714f87629eca6d3%2FScreenshot%202024-07-25%20at%208.39.40PM.png?generation=1721911199454127&alt=media)\nIntuitively this represents the scenario where for any search term, non-relevant documents are never in the index, so rank doesn’t matter since as long as the document has the term, it will always be relevant. As the index size gets smaller compared to the actual global size, this approximation gets better since it becomes increasingly unlikely that any non-relevant document is returned regardless of rank.\n*Interpolated function*:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F65cb63c5452e2bf298f3210ea05702ad%2FScreenshot%202024-07-25%20at%208.40.31PM.png?generation=1721911253003483&alt=media)\nThe alpha parameter represents the contribution of ranked and uniform components to the overall weight. Intuitively this represents the probability of false positives given an index size and global size. If the index and global are the same, then rank weights are directly equivalent to search ranks and the weight of a query term is directly equivalent to the ranks of the neighbour documents with that term. If the index is infinitesimally smaller than global, then rank doesn’t matter and the weight of a query term is only related to the number of neighbour documents with that term.\n\n###### Set cover algorithm\nUsing weights from above, the simple greedy approach is to pick the term that provides the largest increase in total weight in every iteration. I introduce a new parameter, threshold, that represents the ‘saturation’ of weight for an item in the set. For example, if query term a contributes weight 1 to patent x, and threshold is 1, then any query term that also contributes to patent x doesn’t increase the total weight for x, since term a already returns x exactly.\n\nThe algorithm loop is:\n- Try each query term from all 50 neighbours and calculate the increase in weight for all neighbours, subject to threshold\n- Pick the term with highest weight contribution, and remove it from candidates\n- Repeat until maximum weight is reached (all neighbour weights reach threshold) or tokens exceed 50\n\nFor faster computation I represent set cover as a bitset and only save the highest weight per unique bitset during weight calculation\n\n## Parameter optimization\nIn addition to alpha and saturation threshold, I have another threshold for filtering out low weightage terms in the initial calculation. These 3 parameters can be easily tuned, and what I tried with the limited time was Bayesian optimization, treating the full scoring against 1 validation index as 1 black box evaluation. In practice the effectiveness was severely limited due to bad validation indexes, query time limit making evaluations noisy, but I do believe that simple tuning is effective if these are resolved. The initial parameter bounds are:\n- alpha (0, 0.8)\n- weight threshold (0, 0.04)\n- combination threshold (saturation) (0, 0.7)\n\nHere is a 3d visualization (credits: https://github.com/Yamanaka-Lab-TUAT/BOXVIA) of one BO set of trials against a validation index, with EI as acquisition function:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fb590afa1e271e15c3513e295e2f99e09%2FScreenshot%202024-07-25%20at%208.02.05PM.png?generation=1721911655634269&alt=media)\nNote that for this particular run it looks like fully uniform weighting for a range of combination thresholds is not far from the global maxima, and the weight threshold doesn't seem to matter much. \n\n## Closing notes\nTldr my solution consists of a heavy preprocessing pipeline to extract meaningful noun phrases, a greedy set cover algorithm with weighting that tries to approximate the sampling of the validation/test index from the global dataset, and some simple tuning. All the data and code are linked below and I look forward to any comments, feedback, suggestions, especially any mistakes or how I can improve. Thanks!\n\n\nDataset: https://www.kaggle.com/datasets/dilliontan/uspto-preprocessed/\nSubmission notebook: https://www.kaggle.com/code/dilliontan/uspto-optimized-submission\n\n###### Query examples\n\nUS-5516579-A\n23 tokens covers all neighbours\n>detd:\"polymeric-phenolic-esters\" OR detd:\"predominant-dicarboxylic-acid\" OR detd:\"tmac-tma-bpda\" OR ti:\"solid-phase-polycondensation\" OR ab:\"rapid-molding-time\" OR clm:\"particulate-flame-retardants\" OR ab:\"especially-outstanding-toughness\" OR ab:\"ether-glycol-phthalate\" OR ab:\"new-economic-prepolymers\" OR ab:\"either-aliphatic-diacid\" OR clm:\"crystalline-polyester-imide\" OR ab:\"effective-reinforcing-amount\"\n\nUS-8637833-B2\n1 token covers all neighbours\n>detd:\"proton-delivery-efficiency\"\n\nUS-11663219-B1\n1 token covers all neighbours\n>detd:\"editable-buckets\"\n\nUS-8047197-B1\n49 tokens covers 29 neighbours\n>ab:\"tile-roofs\" OR clm:\"pointed-assembly\" OR ab:\"tile-roof\" OR ab:\"coupling-support-structure\" OR detd:\"screw-attached-mount\" OR clm:\"said-primary-barriers\" OR ab:\"phone-stand-assembly\" OR detd:\"vertically-orientated-bracket\" OR ti:\"food-cooling-assembly\" OR ti:\"roof-bracket-apparatus\" OR clm:\"fixedly-arranged-devices\" OR ti:\"rain-monitoring-system\" OR ab:\"tile-support-panels\" OR ab:\"consecutive-inclined-surfaces\" OR ti:\"laptop-support-assembly\" OR clm:\"bent-section-directing\" OR clm:\"convexly-arcuate-prominence\" OR ti:\"rod-holding-system\" OR clm:\"horizontally-orientated-members\" OR clm:\"telecommunications-rack-structure\" OR clm:\"middle-telescoping-member\" OR ti:\"tilt-table-system\" OR detd:\"first-elongate-spaces\" OR ti:\"door-painting-assembly\" OR ab:\"relatively-perpendicular-angle\"\n\nUS-RE30747-E\n49 tokens covers 42 neighbours\n>detd:\"muszik-et\" OR detd:\"structuring-substance\" OR detd:\"hectograph-blankets\" OR detd:\"fill-adhesive-composition\" OR detd:\"pregummed-wall-paper\" OR detd:\"outstanding-writability\" OR detd:\"decalcomania-coating\" OR detd:\"thermo-adhesive-tapes\" OR detd:\"di-dimethylcyclohexyl-adipate\" OR detd:\"single-plasticizers\" OR ti:\"polyvinyl-alcohol-deposits\" OR ti:\"decalcomania-adhesive\" OR ti:\"marine-dye-composition\" OR clm:\"water-dispersive-copolymer\" OR ti:\"pentavalent-vanadium-compound\" OR detd:\"example-starch-hydrate\" OR ab:\"medium-moisture-proof\" OR ti:\"coated-cork-composition\" OR detd:\"least-limitedsolubility\" OR detd:\"serial-tolerance\" OR ab:\"excellent-frangibility\" OR detd:\"difilcultto-remove\" OR detd:\"filly-name\" OR ab:\"stick-application\" OR clm:\"octyl-terephthalic-acid\"\n\nUS-10013212-B2\n49 tokens covers 38 neighbours\n>detd:\"intrusion-detection-database\" OR detd:\"avalon-memory-mapped\" OR clm:\"programmable-hardware-units\" OR ab:\"block-rams\" OR detd:\"reconfigurable-gate-arrays\" OR detd:\"stabilizing-logic-devices\" OR clm:\"programmable-logic-processor\" OR detd:\"partially-reconfigurable-regions\" OR ti:\"heterogeneous-memory-primitives\" OR ab:\"guardband-voltages\" OR ti:\"gpu-virtualisation\" OR ab:\"least-arbitration-rule\" OR ab:\"configuration-programming-block\" OR clm:\"variable-size-ratio\" OR ti:\"immediate-value-network\" OR ab:\"iii-cache-data\" OR ti:\"fast-dynamic-change\" OR ti:\"matrix-storage-method\" OR clm:\"asynchronous-handshaking-scheme\" OR ti:\"hierarchical-sort-acceleration\" OR clm:\"execution-complete-flag\" OR clm:\"coherent-interconnection-protocol\" OR clm:\"intermediate-module-channel\" OR ab:\"correct-storage-address\" OR detd:\"problematic-memory-blocks\"",
    "2940830": "Congratulations on earning a gold medal and achieving 12th place in the competition! @dilliontan \nYour approach to using bigrams and trigrams, along with the detailed explanation of your preprocessing pipeline and weighted set cover algorithm, is impressive. It’s fascinating to see how you tackled the challenge with a combination of meaningful n-grams and advanced techniques like Aho-Corasick for efficient substring matching. Your insights into the nuances of phrase search, particularly in the description field, and the impact of special characters on token savings add valuable depth to your solution. Thanks for sharing your detailed methodology and the links to your dataset and code—these will undoubtedly benefit others looking to refine their own approaches in similar tasks.",
    "2935973": "Thanks for sharing! Would you mind tagging this as a solution writeup so other people can access it from the leaderboard as well?",
    "2936670": ""
  }
}