{
  "id": 522202,
  "title": "6th Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/sash-6th-place-solution",
  "author_name": "",
  "post_date": "2024-08-11T02:57:59.633Z",
  "votes": 24,
  "comment_count": 7,
  "views": 0,
  "content": "<ul>\n<li>Updated on 2024-07-27: I published the notebook at <a href=\"https://www.kaggle.com/code/sash2104/uspto-6th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/sash2104/uspto-6th-place-solution</a></li>\n<li>Updated on 2024-08-11: I published the scripts at <a href=\"https://github.com/sash2104/kaggle-uspto-public\" target=\"_blank\">https://github.com/sash2104/kaggle-uspto-public</a></li>\n<li>Updated on 2024-08-11: 日本語のwriteupを <a href=\"https://github.com/sash2104/kaggle-uspto-public/blob/main/writeup/README_ja.md\" target=\"_blank\">https://github.com/sash2104/kaggle-uspto-public/blob/main/writeup/README_ja.md</a> に追加</li>\n</ul>\n<p>Thanks to the hosts for organizing this competition! I thoroughly enjoyed participating throughout the duration.</p>\n<h2>Summary</h2>\n<p>An example of a query is as follows:<br>\n<code>(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain</code></p>\n<ul>\n<li>Only <code>AND</code> and <code>OR</code> operators are used.<ul>\n<li><code>AND</code> is implied and therefore omitted (<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104</a>)</li></ul></li>\n<li>All fields are utilized in the query.</li>\n<li>The query format is <code>subquery_1 OR subquery_2 OR ... OR subquery_k</code>. The example query consists of four subqueries. The first two subqueries sharing <code>ti:composition</code> to save tokens:<ul>\n<li><code>ti:composition detd:coox detd:pearly</code></li>\n<li><code>ti:composition clm:dye detd:behentrimoinium</code></li>\n<li><code>cpc:A61Q5/12 detd:artichoke detd:genaminox</code></li>\n<li><code>detd:amidoquatsagain</code></li></ul></li>\n<li>For validation, I used a custom searcher implemented in C++ and an alternative metric.</li>\n</ul>\n<h2>Terms</h2>\n<ul>\n<li>Result Set: A list of patents returned from the custom searcher in response to a query. It is not limited to 50 items, and includes all patents that match the query in the index.</li>\n<li>AP'@50: the competition metric, is a flawed implementation of AP@50. (<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981#2791642\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981#2791642</a>)</li>\n<li>Positive: Target 50 patents.</li>\n<li>Negative: All patents that are not Positive. These can be divided into two types based on whether they are definitely included in the test index or not.<ul>\n<li>Certain Negative: Negative patents that exist in test.csv (~12K).</li>\n<li>Potential Negative: Patents that exist in patent_metadata.parquet but not in test.csv (~13M).</li></ul></li>\n<li>Padding: Virtual patents added when the number of results is less than 50.</li>\n</ul>\n<h2>Validation Strategy</h2>\n<p>As discussed in <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169</a> , I considered creating queries using all the patents included in patent_metadata.parquet. However, it seemed impractical to create a Whoosh index that includes all 13 million patents, so I decided to use my own searcher and an alternative metric for validation.</p>\n<p>To conduct the evaluation quickly, I added the following constraints:</p>\n<ul>\n<li>Constraint 1: Queries will not include proximity operators.</li>\n<li>Constraint 2: There will be no weighting of scores among patents that match the query.</li>\n<li>Constraint 3: Queries will not include wildcards.</li>\n</ul>\n<p>Constraint 1 removes the need to consider the positions of words in the patents, Constraint 2 removes the need to consider the frequency of words in the patents, and Constraint 3 removes the need to consider the surface forms of words.</p>\n<p>If only Positives (and Padding) exist in the Result Set, accurate AP'@50 can be calculated even with these constraints. For example, if Positives change from 0, 10, 20, 30, 40, to 50, the AP'@50 changes from 0 to 0.51, 0.76, 0.91, 0.97, and 1.0, respectively.</p>\n<p>When Positives and Negatives are mixed, AP'@50 can only be approximated. For instance, with one Positive and one Negative (and 48 Padding), there are two possible patterns of results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2F4356475ff5ee4752a0f5de78853bbf7a%2Fuspto_ap50_p1n1.png?generation=1721866546870616&amp;alt=media\" alt=\"ap@50\"></p>\n<p>Due to Constraint 2, it is not possible to determine a single pattern for the result, so I assumed that all patterns appear with equal probability. The approximate AP'@50 score in this case is <code>(0.090 + 0.070)/2 = 0.080</code>. Similarly, for n Positives and m Negatives, (n+m)Cm patterns of results are possible. The score for n Positives and m Negatives was taken as the average AP'@50 of 100,000 patterns randomly sampled with replacement from all possible patterns (henceforth referred to as score1(n,m)).</p>\n<p>The graph below shows the plotted scores.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Fabf1280c0f592a0a028e4615b135da90%2Fuspto_score1.png?generation=1721866611625446&amp;alt=media\" alt=\"score1\"></p>\n<p>For example, score1(25,0)=0.842 and score1(50,50)=0.500.</p>\n<p>If only Positives and Certain Negatives are included in the Result Set, <code>score1</code> can be used as is.</p>\n<p>When Potential Negatives are included in the Result Set, the number of Negatives varies probabilistically. Assuming all Potential Negatives are included in the Result Set with probability p, the score can be calculated using the probability mass function (pmf) of the binomial distribution as follows:</p>\n<p>$$<br>\n\\text{score2}(n, m, p) = \\sum_{k=0}^{m} \\binom{m}{k} p^k (1-p)^{m-k} \\text{score1}(n, k) <br>\n$$</p>\n<p>For example, score2(50,50,1.0) = 0.500, score2(50,50,0.1)=0.910, and score2(50,50,0.01)=0.990.</p>\n<p>The validation score was calculated as the average score2(n, m, 0.2) for each of the 2500 rows in the self-created test.csv equivalent. The total number of patents is approximately 13 million, and the number of Potential Negatives that could be included in the test index is around 75,000, so <code>p=75000/13000000 ≒ 0.006</code> seemed appropriate. However, to avoid optimistic validation scores, p=0.2 was used.</p>\n<p>The equivalent of the publication_number in test.csv was prepared as follows:</p>\n<ul>\n<li>2500 patents since 1975 and <code>len(cpc_codes) &gt; 0</code></li>\n</ul>\n<p>When the test index was created by simply filtering patents after 1975, the validation score and LB score seemed slightly off. Analyzing the published train index and LB revealed that design patents starting with US-D were rarely included in the train index and test.csv used for LB calculation. Upon investigating several design patents, I found that cpc_codes were often empty, so I added the filter <code>len(cpc_codes) &gt; 0</code>. I also confirmed that patents with empty cpc_codes were not included in the publication_number column of LB's test.csv.</p>\n<p>With the above modifications, the difference between the LB score and the validation score was generally within 0.01.</p>\n<h2>Solution</h2>\n<ul>\n<li>Phase 1: Find subquery candidates</li>\n<li>Phase 2: Generate queries using beam search to maximize <code>score2(n, m, 0.2)</code></li>\n</ul>\n<p>To achieve fast execution speed, as with my own searcher, everything was implemented in C++.</p>\n<h3>Phase 1: Finding Subquery Candidates</h3>\n<ul>\n<li>n-shot subquery<ul>\n<li>A subquery that narrows down to a specific single patent. For example, ti:margaric narrows down to US-2023092870-A1 because it is the only patent containing margaric in the title. Similarly, the cpc_codes A46B13/005 and A47L7/0061 co-occur only in US-2023025335-A1, so <code>cpc:A46B13/005 AND cpc:A47L7/0061</code> narrows down to US-2023025335-A1.</li>\n<li>Pre-calculated and embedded.</li>\n<li>Found that 5.9M patents could be narrowed down to one with a single token, and an additional 6.7M patents could be narrowed down with two tokens.</li></ul></li>\n<li>conjunctive subquery<ul>\n<li>A subquery that includes multiple Positives, no Certain Negatives, and at most <code>l</code> Potential Negatives.</li>\n<li>Combination of 2-3 tokens using AND. </li>\n<li>Explored combinations of cpc_codes, words in title and abstract (frequency &lt;=400,000), claims (frequency &lt;=100,000), and description (frequency &lt;=10,000) using DFS (Depth-First Search).</li>\n<li>Explored most cases fully, but some were time-limited.</li>\n<li>Avoided Certain Negatives as they significantly reduce the score and facilitate efficient DFS pruning.</li>\n<li>Conducted beam search with <code>l=0</code> and <code>l=1</code> for Phase 2. This improved validation score by about 0.003.</li>\n<li>Calculated on-the-fly rather than pre-calculating.</li></ul></li>\n</ul>\n<h3>Phase 2: Beam Search</h3>\n<p>The evaluation metric used was score2(n, m, 0.2), and I performed beam search with a width W for each number of tokens used. For the submission, I set W=100. Increasing W slightly improved the validation score, but even increasing W by a factor of 10 only changed the validation score by about 0.001. Additionally, if subqueries contained common tokens, they were combined to save on the number of tokens used.</p>\n<p>For example, the query for US-7507696-B2 is as follows:  <br>\n<code>(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain</code></p>\n<p>This query matches only the 50 target patents of US-7507696-B2 out of approximately 13 million patents. The query is broken down into four subqueries:</p>\n<ul>\n<li><code>ti:composition detd:coox detd:pearly</code></li>\n<li><code>ti:composition clm:dye detd:behentrimoinium</code></li>\n<li><code>cpc:A61Q5/12 detd:artichoke detd:genaminox</code></li>\n<li><code>detd:amidoquatsagain</code></li>\n</ul>\n<p>The first three subqueries are conjunctive subqueries, and the first and second share the common token ti:composition, so they are combined. The final subquery, detd:amidoquatsagain, is an n-shot subquery and matches only US-8044007-B2.</p>\n<p>Like the example query shown, the proportion of perfect queries generated that matched only the 50 target patents out of all patents was about 6%.</p>\n<p>The distribution of score2 in the submission with the best private score is as follows.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Ff47a9d31197257e42b2dd840194ebaaa%2Fscore2_distribution.png?generation=1721866520258494&amp;alt=media\" alt=\"score2_distribution\"></p>\n<h2>What wasn't used/worked</h2>\n<ul>\n<li><code>NOT</code><ul>\n<li>I thought it might be possible to improve the score in the form of <code>(subquery_a OR ...) NOT (subquery_b OR ...)</code> by roughly narrowing down the patents and then eliminating the Negatives in the subqueries after NOT, but I didn't have time to implement and test it.</li></ul></li>\n<li>Other metaheuristics in Phase 2<ul>\n<li>I tried replacing candidates using hill climbing or simulated annealing, but simply replacing candidates randomly was inferior to beam search. Moreover, even if I worked hard to improve Phase 2 alone, the validation score seemed to improve by less than 0.01, so I didn't delve too deeply into it.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "2935060",
      "postDate": "07/25/2024 00:17:55",
      "content": "<ul>\n<li>Updated on 2024-07-27: I published the notebook at <a href=\"https://www.kaggle.com/code/sash2104/uspto-6th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/sash2104/uspto-6th-place-solution</a></li>\n<li>Updated on 2024-08-11: I published the scripts at <a href=\"https://github.com/sash2104/kaggle-uspto-public\" target=\"_blank\">https://github.com/sash2104/kaggle-uspto-public</a></li>\n<li>Updated on 2024-08-11: 日本語のwriteupを <a href=\"https://github.com/sash2104/kaggle-uspto-public/blob/main/writeup/README_ja.md\" target=\"_blank\">https://github.com/sash2104/kaggle-uspto-public/blob/main/writeup/README_ja.md</a> に追加</li>\n</ul>\n<p>Thanks to the hosts for organizing this competition! I thoroughly enjoyed participating throughout the duration.</p>\n<h2>Summary</h2>\n<p>An example of a query is as follows:<br>\n<code>(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain</code></p>\n<ul>\n<li>Only <code>AND</code> and <code>OR</code> operators are used.<ul>\n<li><code>AND</code> is implied and therefore omitted (<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104</a>)</li></ul></li>\n<li>All fields are utilized in the query.</li>\n<li>The query format is <code>subquery_1 OR subquery_2 OR ... OR subquery_k</code>. The example query consists of four subqueries. The first two subqueries sharing <code>ti:composition</code> to save tokens:<ul>\n<li><code>ti:composition detd:coox detd:pearly</code></li>\n<li><code>ti:composition clm:dye detd:behentrimoinium</code></li>\n<li><code>cpc:A61Q5/12 detd:artichoke detd:genaminox</code></li>\n<li><code>detd:amidoquatsagain</code></li></ul></li>\n<li>For validation, I used a custom searcher implemented in C++ and an alternative metric.</li>\n</ul>\n<h2>Terms</h2>\n<ul>\n<li>Result Set: A list of patents returned from the custom searcher in response to a query. It is not limited to 50 items, and includes all patents that match the query in the index.</li>\n<li>AP'@50: the competition metric, is a flawed implementation of AP@50. (<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981#2791642\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981#2791642</a>)</li>\n<li>Positive: Target 50 patents.</li>\n<li>Negative: All patents that are not Positive. These can be divided into two types based on whether they are definitely included in the test index or not.<ul>\n<li>Certain Negative: Negative patents that exist in test.csv (~12K).</li>\n<li>Potential Negative: Patents that exist in patent_metadata.parquet but not in test.csv (~13M).</li></ul></li>\n<li>Padding: Virtual patents added when the number of results is less than 50.</li>\n</ul>\n<h2>Validation Strategy</h2>\n<p>As discussed in <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169</a> , I considered creating queries using all the patents included in patent_metadata.parquet. However, it seemed impractical to create a Whoosh index that includes all 13 million patents, so I decided to use my own searcher and an alternative metric for validation.</p>\n<p>To conduct the evaluation quickly, I added the following constraints:</p>\n<ul>\n<li>Constraint 1: Queries will not include proximity operators.</li>\n<li>Constraint 2: There will be no weighting of scores among patents that match the query.</li>\n<li>Constraint 3: Queries will not include wildcards.</li>\n</ul>\n<p>Constraint 1 removes the need to consider the positions of words in the patents, Constraint 2 removes the need to consider the frequency of words in the patents, and Constraint 3 removes the need to consider the surface forms of words.</p>\n<p>If only Positives (and Padding) exist in the Result Set, accurate AP'@50 can be calculated even with these constraints. For example, if Positives change from 0, 10, 20, 30, 40, to 50, the AP'@50 changes from 0 to 0.51, 0.76, 0.91, 0.97, and 1.0, respectively.</p>\n<p>When Positives and Negatives are mixed, AP'@50 can only be approximated. For instance, with one Positive and one Negative (and 48 Padding), there are two possible patterns of results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2F4356475ff5ee4752a0f5de78853bbf7a%2Fuspto_ap50_p1n1.png?generation=1721866546870616&amp;alt=media\" alt=\"ap@50\"></p>\n<p>Due to Constraint 2, it is not possible to determine a single pattern for the result, so I assumed that all patterns appear with equal probability. The approximate AP'@50 score in this case is <code>(0.090 + 0.070)/2 = 0.080</code>. Similarly, for n Positives and m Negatives, (n+m)Cm patterns of results are possible. The score for n Positives and m Negatives was taken as the average AP'@50 of 100,000 patterns randomly sampled with replacement from all possible patterns (henceforth referred to as score1(n,m)).</p>\n<p>The graph below shows the plotted scores.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Fabf1280c0f592a0a028e4615b135da90%2Fuspto_score1.png?generation=1721866611625446&amp;alt=media\" alt=\"score1\"></p>\n<p>For example, score1(25,0)=0.842 and score1(50,50)=0.500.</p>\n<p>If only Positives and Certain Negatives are included in the Result Set, <code>score1</code> can be used as is.</p>\n<p>When Potential Negatives are included in the Result Set, the number of Negatives varies probabilistically. Assuming all Potential Negatives are included in the Result Set with probability p, the score can be calculated using the probability mass function (pmf) of the binomial distribution as follows:</p>\n<p>$$<br>\n\\text{score2}(n, m, p) = \\sum_{k=0}^{m} \\binom{m}{k} p^k (1-p)^{m-k} \\text{score1}(n, k) <br>\n$$</p>\n<p>For example, score2(50,50,1.0) = 0.500, score2(50,50,0.1)=0.910, and score2(50,50,0.01)=0.990.</p>\n<p>The validation score was calculated as the average score2(n, m, 0.2) for each of the 2500 rows in the self-created test.csv equivalent. The total number of patents is approximately 13 million, and the number of Potential Negatives that could be included in the test index is around 75,000, so <code>p=75000/13000000 ≒ 0.006</code> seemed appropriate. However, to avoid optimistic validation scores, p=0.2 was used.</p>\n<p>The equivalent of the publication_number in test.csv was prepared as follows:</p>\n<ul>\n<li>2500 patents since 1975 and <code>len(cpc_codes) &gt; 0</code></li>\n</ul>\n<p>When the test index was created by simply filtering patents after 1975, the validation score and LB score seemed slightly off. Analyzing the published train index and LB revealed that design patents starting with US-D were rarely included in the train index and test.csv used for LB calculation. Upon investigating several design patents, I found that cpc_codes were often empty, so I added the filter <code>len(cpc_codes) &gt; 0</code>. I also confirmed that patents with empty cpc_codes were not included in the publication_number column of LB's test.csv.</p>\n<p>With the above modifications, the difference between the LB score and the validation score was generally within 0.01.</p>\n<h2>Solution</h2>\n<ul>\n<li>Phase 1: Find subquery candidates</li>\n<li>Phase 2: Generate queries using beam search to maximize <code>score2(n, m, 0.2)</code></li>\n</ul>\n<p>To achieve fast execution speed, as with my own searcher, everything was implemented in C++.</p>\n<h3>Phase 1: Finding Subquery Candidates</h3>\n<ul>\n<li>n-shot subquery<ul>\n<li>A subquery that narrows down to a specific single patent. For example, ti:margaric narrows down to US-2023092870-A1 because it is the only patent containing margaric in the title. Similarly, the cpc_codes A46B13/005 and A47L7/0061 co-occur only in US-2023025335-A1, so <code>cpc:A46B13/005 AND cpc:A47L7/0061</code> narrows down to US-2023025335-A1.</li>\n<li>Pre-calculated and embedded.</li>\n<li>Found that 5.9M patents could be narrowed down to one with a single token, and an additional 6.7M patents could be narrowed down with two tokens.</li></ul></li>\n<li>conjunctive subquery<ul>\n<li>A subquery that includes multiple Positives, no Certain Negatives, and at most <code>l</code> Potential Negatives.</li>\n<li>Combination of 2-3 tokens using AND. </li>\n<li>Explored combinations of cpc_codes, words in title and abstract (frequency &lt;=400,000), claims (frequency &lt;=100,000), and description (frequency &lt;=10,000) using DFS (Depth-First Search).</li>\n<li>Explored most cases fully, but some were time-limited.</li>\n<li>Avoided Certain Negatives as they significantly reduce the score and facilitate efficient DFS pruning.</li>\n<li>Conducted beam search with <code>l=0</code> and <code>l=1</code> for Phase 2. This improved validation score by about 0.003.</li>\n<li>Calculated on-the-fly rather than pre-calculating.</li></ul></li>\n</ul>\n<h3>Phase 2: Beam Search</h3>\n<p>The evaluation metric used was score2(n, m, 0.2), and I performed beam search with a width W for each number of tokens used. For the submission, I set W=100. Increasing W slightly improved the validation score, but even increasing W by a factor of 10 only changed the validation score by about 0.001. Additionally, if subqueries contained common tokens, they were combined to save on the number of tokens used.</p>\n<p>For example, the query for US-7507696-B2 is as follows:  <br>\n<code>(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain</code></p>\n<p>This query matches only the 50 target patents of US-7507696-B2 out of approximately 13 million patents. The query is broken down into four subqueries:</p>\n<ul>\n<li><code>ti:composition detd:coox detd:pearly</code></li>\n<li><code>ti:composition clm:dye detd:behentrimoinium</code></li>\n<li><code>cpc:A61Q5/12 detd:artichoke detd:genaminox</code></li>\n<li><code>detd:amidoquatsagain</code></li>\n</ul>\n<p>The first three subqueries are conjunctive subqueries, and the first and second share the common token ti:composition, so they are combined. The final subquery, detd:amidoquatsagain, is an n-shot subquery and matches only US-8044007-B2.</p>\n<p>Like the example query shown, the proportion of perfect queries generated that matched only the 50 target patents out of all patents was about 6%.</p>\n<p>The distribution of score2 in the submission with the best private score is as follows.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Ff47a9d31197257e42b2dd840194ebaaa%2Fscore2_distribution.png?generation=1721866520258494&amp;alt=media\" alt=\"score2_distribution\"></p>\n<h2>What wasn't used/worked</h2>\n<ul>\n<li><code>NOT</code><ul>\n<li>I thought it might be possible to improve the score in the form of <code>(subquery_a OR ...) NOT (subquery_b OR ...)</code> by roughly narrowing down the patents and then eliminating the Negatives in the subqueries after NOT, but I didn't have time to implement and test it.</li></ul></li>\n<li>Other metaheuristics in Phase 2<ul>\n<li>I tried replacing candidates using hill climbing or simulated annealing, but simply replacing candidates randomly was inferior to beam search. Moreover, even if I worked hard to improve Phase 2 alone, the validation score seemed to improve by less than 0.01, so I didn't delve too deeply into it.</li></ul></li>\n</ul>",
      "rawMarkdown": "Updated on 2024-07-27: I published the notebook at https://www.kaggle.com/code/sash2104/uspto-6th-place-solution\n- Updated on 2024-08-11: I published the scripts at https://github.com/sash2104/kaggle-uspto-public\n- Updated on 2024-08-11: 日本語のwriteupを https://github.com/sash2104/kaggle-uspto-public/blob/main/writeup/README_ja.md に追加\n\nThanks to the hosts for organizing this competition! I thoroughly enjoyed participating throughout the duration.\n\n## Summary\nAn example of a query is as follows:\n`(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain`\n- Only `AND` and `OR` operators are used.\n  - `AND` is implied and therefore omitted (https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104)\n- All fields are utilized in the query.\n- The query format is `subquery_1 OR subquery_2 OR ... OR subquery_k`. The example query consists of four subqueries. The first two subqueries sharing `ti:composition` to save tokens:\n  - `ti:composition detd:coox detd:pearly`\n  - `ti:composition clm:dye detd:behentrimoinium`\n  - `cpc:A61Q5/12 detd:artichoke detd:genaminox`\n  - `detd:amidoquatsagain`\n- For validation, I used a custom searcher implemented in C++ and an alternative metric.\n\n## Terms\n- Result Set: A list of patents returned from the custom searcher in response to a query. It is not limited to 50 items, and includes all patents that match the query in the index.\n- AP'@50: the competition metric, is a flawed implementation of AP@50. (https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981#2791642)\n- Positive: Target 50 patents.\n- Negative: All patents that are not Positive. These can be divided into two types based on whether they are definitely included in the test index or not.\n  - Certain Negative: Negative patents that exist in test.csv (~12K).\n  - Potential Negative: Patents that exist in patent_metadata.parquet but not in test.csv (~13M).\n- Padding: Virtual patents added when the number of results is less than 50.\n\n## Validation Strategy\nAs discussed in https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169 , I considered creating queries using all the patents included in patent_metadata.parquet. However, it seemed impractical to create a Whoosh index that includes all 13 million patents, so I decided to use my own searcher and an alternative metric for validation.\n\nTo conduct the evaluation quickly, I added the following constraints:\n- Constraint 1: Queries will not include proximity operators.\n- Constraint 2: There will be no weighting of scores among patents that match the query.\n- Constraint 3: Queries will not include wildcards.\n\nConstraint 1 removes the need to consider the positions of words in the patents, Constraint 2 removes the need to consider the frequency of words in the patents, and Constraint 3 removes the need to consider the surface forms of words.\n\nIf only Positives (and Padding) exist in the Result Set, accurate AP'@50 can be calculated even with these constraints. For example, if Positives change from 0, 10, 20, 30, 40, to 50, the AP'@50 changes from 0 to 0.51, 0.76, 0.91, 0.97, and 1.0, respectively.\n\nWhen Positives and Negatives are mixed, AP'@50 can only be approximated. For instance, with one Positive and one Negative (and 48 Padding), there are two possible patterns of results.\n\n![ap@50](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2F4356475ff5ee4752a0f5de78853bbf7a%2Fuspto_ap50_p1n1.png?generation=1721866546870616&alt=media)\n\n\nDue to Constraint 2, it is not possible to determine a single pattern for the result, so I assumed that all patterns appear with equal probability. The approximate AP'@50 score in this case is `(0.090 + 0.070)/2 = 0.080`. Similarly, for n Positives and m Negatives, (n+m)Cm patterns of results are possible. The score for n Positives and m Negatives was taken as the average AP'@50 of 100,000 patterns randomly sampled with replacement from all possible patterns (henceforth referred to as score1(n,m)).\n\nThe graph below shows the plotted scores.\n\n![score1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Fabf1280c0f592a0a028e4615b135da90%2Fuspto_score1.png?generation=1721866611625446&alt=media)\n\nFor example, score1(25,0)=0.842 and score1(50,50)=0.500.\n\nIf only Positives and Certain Negatives are included in the Result Set, `score1` can be used as is.\n\nWhen Potential Negatives are included in the Result Set, the number of Negatives varies probabilistically. Assuming all Potential Negatives are included in the Result Set with probability p, the score can be calculated using the probability mass function (pmf) of the binomial distribution as follows:\n\n$$\n\\text{score2}(n, m, p) = \\sum_{k=0}^{m} \\binom{m}{k} p^k (1-p)^{m-k} \\text{score1}(n, k) \n$$\n\nFor example, score2(50,50,1.0) = 0.500, score2(50,50,0.1)=0.910, and score2(50,50,0.01)=0.990.\n\nThe validation score was calculated as the average score2(n, m, 0.2) for each of the 2500 rows in the self-created test.csv equivalent. The total number of patents is approximately 13 million, and the number of Potential Negatives that could be included in the test index is around 75,000, so `p=75000/13000000 ≒ 0.006` seemed appropriate. However, to avoid optimistic validation scores, p=0.2 was used.\n\nThe equivalent of the publication_number in test.csv was prepared as follows:\n- 2500 patents since 1975 and `len(cpc_codes) > 0`\n\nWhen the test index was created by simply filtering patents after 1975, the validation score and LB score seemed slightly off. Analyzing the published train index and LB revealed that design patents starting with US-D were rarely included in the train index and test.csv used for LB calculation. Upon investigating several design patents, I found that cpc_codes were often empty, so I added the filter `len(cpc_codes) > 0`. I also confirmed that patents with empty cpc_codes were not included in the publication_number column of LB's test.csv.\n\nWith the above modifications, the difference between the LB score and the validation score was generally within 0.01.\n\n## Solution\n- Phase 1: Find subquery candidates\n- Phase 2: Generate queries using beam search to maximize `score2(n, m, 0.2)`\n\nTo achieve fast execution speed, as with my own searcher, everything was implemented in C++.\n\n### Phase 1: Finding Subquery Candidates\n- n-shot subquery\n  - A subquery that narrows down to a specific single patent. For example, ti:margaric narrows down to US-2023092870-A1 because it is the only patent containing margaric in the title. Similarly, the cpc_codes A46B13/005 and A47L7/0061 co-occur only in US-2023025335-A1, so `cpc:A46B13/005 AND cpc:A47L7/0061` narrows down to US-2023025335-A1.\n  - Pre-calculated and embedded.\n    - Found that 5.9M patents could be narrowed down to one with a single token, and an additional 6.7M patents could be narrowed down with two tokens.\n- conjunctive subquery\n  - A subquery that includes multiple Positives, no Certain Negatives, and at most `l` Potential Negatives.\n  - Combination of 2-3 tokens using AND. \n    - Explored combinations of cpc_codes, words in title and abstract (frequency <=400,000), claims (frequency <=100,000), and description (frequency <=10,000) using DFS (Depth-First Search).\n    - Explored most cases fully, but some were time-limited.\n  - Avoided Certain Negatives as they significantly reduce the score and facilitate efficient DFS pruning.\n  - Conducted beam search with `l=0` and `l=1` for Phase 2. This improved validation score by about 0.003.\n  - Calculated on-the-fly rather than pre-calculating.\n\n### Phase 2: Beam Search\nThe evaluation metric used was score2(n, m, 0.2), and I performed beam search with a width W for each number of tokens used. For the submission, I set W=100. Increasing W slightly improved the validation score, but even increasing W by a factor of 10 only changed the validation score by about 0.001. Additionally, if subqueries contained common tokens, they were combined to save on the number of tokens used.\n\nFor example, the query for US-7507696-B2 is as follows:  \n`(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain`\n\nThis query matches only the 50 target patents of US-7507696-B2 out of approximately 13 million patents. The query is broken down into four subqueries:\n\n- `ti:composition detd:coox detd:pearly`\n- `ti:composition clm:dye detd:behentrimoinium`\n- `cpc:A61Q5/12 detd:artichoke detd:genaminox`\n- `detd:amidoquatsagain`\n\nThe first three subqueries are conjunctive subqueries, and the first and second share the common token ti:composition, so they are combined. The final subquery, detd:amidoquatsagain, is an n-shot subquery and matches only US-8044007-B2.\n\nLike the example query shown, the proportion of perfect queries generated that matched only the 50 target patents out of all patents was about 6%.\n\nThe distribution of score2 in the submission with the best private score is as follows.\n\n![score2_distribution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Ff47a9d31197257e42b2dd840194ebaaa%2Fscore2_distribution.png?generation=1721866520258494&alt=media)\n\n## What wasn't used/worked\n- `NOT`\n  - I thought it might be possible to improve the score in the form of `(subquery_a OR ...) NOT (subquery_b OR ...)` by roughly narrowing down the patents and then eliminating the Negatives in the subqueries after NOT, but I didn't have time to implement and test it.\n- Other metaheuristics in Phase 2\n  - I tried replacing candidates using hill climbing or simulated annealing, but simply replacing candidates randomly was inferior to beam search. Moreover, even if I worked hard to improve Phase 2 alone, the validation score seemed to improve by less than 0.01, so I didn't delve too deeply into it.",
      "votes": null
    },
    {
      "id": "2935303",
      "postDate": "07/25/2024 06:36:17",
      "content": "<p>Congratulations on winning the 6th place in this competition. Thanks for sharing your insights. </p>",
      "rawMarkdown": "Congratulations on winning the 6th place in this competition. Thanks for sharing your insights.",
      "votes": null
    },
    {
      "id": "2936219",
      "postDate": "07/25/2024 23:18:20",
      "content": "<p>Thanks for an amazing solution with a custom C++ code, and maybe the 1st without the magic!</p>\n<p>Does your custom searcher search the whole patent data parquet in the disk <code>kaggle/input</code>, or search only partial negative candidates on RAM? Can I ask what data are on RAM?</p>\n<p>Congratulations for the solo gold!</p>",
      "rawMarkdown": "Thanks for an amazing solution with a custom C++ code, and maybe the 1st without the magic!\n\nDoes your custom searcher search the whole patent data parquet in the disk `kaggle/input`, or search only partial negative candidates on RAM? Can I ask what data are on RAM?\n\nCongratulations for the solo gold!",
      "votes": null
    },
    {
      "id": "2936890",
      "postDate": "07/26/2024 14:22:14",
      "content": "<p>Thank you!</p>\n<p>The data after performing similar preprocessing as in the 1st/2nd place solutions is loaded into RAM.</p>\n<ul>\n<li>The search Index contains the whole patent in patent_metadata.parquet.</li>\n<li>Convert patents and words into IDs instead of using strings.</li>\n<li>Create a two-dimensional variable-length list from patents to words and from words to patents.<ul>\n<li>In the list from patents to words, exclude patents and words that do not appear in the patents included in test.csv.</li>\n<li>Similarly, in the list from words to patents, exclude words that do not appear in those patents.</li></ul></li>\n<li>And so on.</li>\n</ul>\n<p>For explanatory purposes, I have separated the searcher and solver, but in reality, they are implemented in a single C++ file and are tightly coupled. The preprocessing scripts and the notebook will be organized and made public later, so please refer to those for details.</p>",
      "rawMarkdown": "Thank you!\n\nThe data after performing similar preprocessing as in the 1st/2nd place solutions is loaded into RAM.\n- The search Index contains the whole patent in patent_metadata.parquet.\n- Convert patents and words into IDs instead of using strings.\n- Create a two-dimensional variable-length list from patents to words and from words to patents.\n  - In the list from patents to words, exclude patents and words that do not appear in the patents included in test.csv.\n  - Similarly, in the list from words to patents, exclude words that do not appear in those patents.\n- And so on.\n\nFor explanatory purposes, I have separated the searcher and solver, but in reality, they are implemented in a single C++ file and are tightly coupled. The preprocessing scripts and the notebook will be organized and made public later, so please refer to those for details.",
      "votes": null
    },
    {
      "id": "2937242",
      "postDate": "07/26/2024 21:24:56",
      "content": "<p>Thanks! I thought accessing the disk multiple times is too slow, but handling the whole patent was also impossible in terms of time and RAM. Amazing! Congratulations again!</p>",
      "rawMarkdown": "Thanks! I thought accessing the disk multiple times is too slow, but handling the whole patent was also impossible in terms of time and RAM. Amazing! Congratulations again!",
      "votes": null
    },
    {
      "id": "2939851",
      "postDate": "07/29/2024 14:54:54",
      "content": "<p>Nice solution, very similar to mine ! But much more efficient with C++… So you had access to more vocabulary without raising memory errors. Very clean and impressing !</p>\n<p>A well deserved cahsprize 🥳</p>",
      "rawMarkdown": "Nice solution, very similar to mine ! But much more efficient with C++... So you had access to more vocabulary without raising memory errors. Very clean and impressing !\n\nA well deserved cahsprize 🥳",
      "votes": null
    },
    {
      "id": "2940228",
      "postDate": "07/29/2024 21:47:24",
      "content": "<p>Thanks for sharing your solution and congrats on the the prize. </p>\n<p>There is one thing that i don't understand: what is the purpose of distinguish certain negative and potential negative? if it is solely related to how the MAP@50 is calculated, why avoiding certain negative can increase the score? thanks!</p>",
      "rawMarkdown": "Thanks for sharing your solution and congrats on the the prize. \n\nThere is one thing that i don't understand: what is the purpose of distinguish certain negative and potential negative? if it is solely related to how the MAP@50 is calculated, why avoiding certain negative can increase the score? thanks!",
      "votes": null
    },
    {
      "id": "2940809",
      "postDate": "07/30/2024 13:20:44",
      "content": "<p>Thanks!</p>\n<p>There are two reasons:</p>\n<ul>\n<li>If Certain Negatives are included, the score on the LB decreases. For example, Consider a <code>query_a</code> with 20 Positives + 20 Certain Negatives and a <code>query_b</code> with 20 Positives + 20 Potential Negatives on the validation index that includes all patents. In this case, the expected value of AP'@50 (≒LB score) for <code>query_a</code> is score1(20,20)=0.489. On the other hand, since Potential Negatives are only probabilistically included in the test index, the expected value of AP'@50 for <code>query_b</code> will be much higher than 0.489.</li>\n<li>The calculation efficiency of DFS for conjunctive subqueries in Phase 1 improves. The condition of not including Certain Negatives can be confirmed by checking only Certain Negatives, which are included in about 1% of all patents, making it an effective pruning.</li>\n</ul>",
      "rawMarkdown": "Thanks!\n\nThere are two reasons:\n- If Certain Negatives are included, the score on the LB decreases. For example, Consider a `query_a` with 20 Positives + 20 Certain Negatives and a `query_b` with 20 Positives + 20 Potential Negatives on the validation index that includes all patents. In this case, the expected value of AP'@50 (≒LB score) for `query_a` is score1(20,20)=0.489. On the other hand, since Potential Negatives are only probabilistically included in the test index, the expected value of AP'@50 for `query_b` will be much higher than 0.489.\n- The calculation efficiency of DFS for conjunctive subqueries in Phase 1 improves. The condition of not including Certain Negatives can be confirmed by checking only Certain Negatives, which are included in about 1% of all patents, making it an effective pruning.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2935303,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/25/2024 06:36:17",
      "content": "<p>Congratulations on winning the 6th place in this competition. Thanks for sharing your insights. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2936219,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "07/25/2024 23:18:20",
      "content": "<p>Thanks for an amazing solution with a custom C++ code, and maybe the 1st without the magic!</p>\n<p>Does your custom searcher search the whole patent data parquet in the disk <code>kaggle/input</code>, or search only partial negative candidates on RAM? Can I ask what data are on RAM?</p>\n<p>Congratulations for the solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2936890,
          "author_name": "sash2104",
          "author_url": "",
          "post_date": "07/26/2024 14:22:14",
          "content": "<p>Thank you!</p>\n<p>The data after performing similar preprocessing as in the 1st/2nd place solutions is loaded into RAM.</p>\n<ul>\n<li>The search Index contains the whole patent in patent_metadata.parquet.</li>\n<li>Convert patents and words into IDs instead of using strings.</li>\n<li>Create a two-dimensional variable-length list from patents to words and from words to patents.<ul>\n<li>In the list from patents to words, exclude patents and words that do not appear in the patents included in test.csv.</li>\n<li>Similarly, in the list from words to patents, exclude words that do not appear in those patents.</li></ul></li>\n<li>And so on.</li>\n</ul>\n<p>For explanatory purposes, I have separated the searcher and solver, but in reality, they are implemented in a single C++ file and are tightly coupled. The preprocessing scripts and the notebook will be organized and made public later, so please refer to those for details.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2937242,
              "author_name": "junkoda",
              "author_url": "",
              "post_date": "07/26/2024 21:24:56",
              "content": "<p>Thanks! I thought accessing the disk multiple times is too slow, but handling the whole patent was also impossible in terms of time and RAM. Amazing! Congratulations again!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2939851,
      "author_name": "vincentschuler",
      "author_url": "",
      "post_date": "07/29/2024 14:54:54",
      "content": "<p>Nice solution, very similar to mine ! But much more efficient with C++… So you had access to more vocabulary without raising memory errors. Very clean and impressing !</p>\n<p>A well deserved cahsprize 🥳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2940228,
      "author_name": "frank6526",
      "author_url": "",
      "post_date": "07/29/2024 21:47:24",
      "content": "<p>Thanks for sharing your solution and congrats on the the prize. </p>\n<p>There is one thing that i don't understand: what is the purpose of distinguish certain negative and potential negative? if it is solely related to how the MAP@50 is calculated, why avoiding certain negative can increase the score? thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2940809,
          "author_name": "sash2104",
          "author_url": "",
          "post_date": "07/30/2024 13:20:44",
          "content": "<p>Thanks!</p>\n<p>There are two reasons:</p>\n<ul>\n<li>If Certain Negatives are included, the score on the LB decreases. For example, Consider a <code>query_a</code> with 20 Positives + 20 Certain Negatives and a <code>query_b</code> with 20 Positives + 20 Potential Negatives on the validation index that includes all patents. In this case, the expected value of AP'@50 (≒LB score) for <code>query_a</code> is score1(20,20)=0.489. On the other hand, since Potential Negatives are only probabilistically included in the test index, the expected value of AP'@50 for <code>query_b</code> will be much higher than 0.489.</li>\n<li>The calculation efficiency of DFS for conjunctive subqueries in Phase 1 improves. The condition of not including Certain Negatives can be confirmed by checking only Certain Negatives, which are included in about 1% of all patents, making it an effective pruning.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2935060": "Updated on 2024-07-27: I published the notebook at https://www.kaggle.com/code/sash2104/uspto-6th-place-solution\n- Updated on 2024-08-11: I published the scripts at https://github.com/sash2104/kaggle-uspto-public\n- Updated on 2024-08-11: 日本語のwriteupを https://github.com/sash2104/kaggle-uspto-public/blob/main/writeup/README_ja.md に追加\n\nThanks to the hosts for organizing this competition! I thoroughly enjoyed participating throughout the duration.\n\n## Summary\nAn example of a query is as follows:\n`(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain`\n- Only `AND` and `OR` operators are used.\n  - `AND` is implied and therefore omitted (https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104)\n- All fields are utilized in the query.\n- The query format is `subquery_1 OR subquery_2 OR ... OR subquery_k`. The example query consists of four subqueries. The first two subqueries sharing `ti:composition` to save tokens:\n  - `ti:composition detd:coox detd:pearly`\n  - `ti:composition clm:dye detd:behentrimoinium`\n  - `cpc:A61Q5/12 detd:artichoke detd:genaminox`\n  - `detd:amidoquatsagain`\n- For validation, I used a custom searcher implemented in C++ and an alternative metric.\n\n## Terms\n- Result Set: A list of patents returned from the custom searcher in response to a query. It is not limited to 50 items, and includes all patents that match the query in the index.\n- AP'@50: the competition metric, is a flawed implementation of AP@50. (https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981#2791642)\n- Positive: Target 50 patents.\n- Negative: All patents that are not Positive. These can be divided into two types based on whether they are definitely included in the test index or not.\n  - Certain Negative: Negative patents that exist in test.csv (~12K).\n  - Potential Negative: Patents that exist in patent_metadata.parquet but not in test.csv (~13M).\n- Padding: Virtual patents added when the number of results is less than 50.\n\n## Validation Strategy\nAs discussed in https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169 , I considered creating queries using all the patents included in patent_metadata.parquet. However, it seemed impractical to create a Whoosh index that includes all 13 million patents, so I decided to use my own searcher and an alternative metric for validation.\n\nTo conduct the evaluation quickly, I added the following constraints:\n- Constraint 1: Queries will not include proximity operators.\n- Constraint 2: There will be no weighting of scores among patents that match the query.\n- Constraint 3: Queries will not include wildcards.\n\nConstraint 1 removes the need to consider the positions of words in the patents, Constraint 2 removes the need to consider the frequency of words in the patents, and Constraint 3 removes the need to consider the surface forms of words.\n\nIf only Positives (and Padding) exist in the Result Set, accurate AP'@50 can be calculated even with these constraints. For example, if Positives change from 0, 10, 20, 30, 40, to 50, the AP'@50 changes from 0 to 0.51, 0.76, 0.91, 0.97, and 1.0, respectively.\n\nWhen Positives and Negatives are mixed, AP'@50 can only be approximated. For instance, with one Positive and one Negative (and 48 Padding), there are two possible patterns of results.\n\n![ap@50](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2F4356475ff5ee4752a0f5de78853bbf7a%2Fuspto_ap50_p1n1.png?generation=1721866546870616&alt=media)\n\n\nDue to Constraint 2, it is not possible to determine a single pattern for the result, so I assumed that all patterns appear with equal probability. The approximate AP'@50 score in this case is `(0.090 + 0.070)/2 = 0.080`. Similarly, for n Positives and m Negatives, (n+m)Cm patterns of results are possible. The score for n Positives and m Negatives was taken as the average AP'@50 of 100,000 patterns randomly sampled with replacement from all possible patterns (henceforth referred to as score1(n,m)).\n\nThe graph below shows the plotted scores.\n\n![score1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Fabf1280c0f592a0a028e4615b135da90%2Fuspto_score1.png?generation=1721866611625446&alt=media)\n\nFor example, score1(25,0)=0.842 and score1(50,50)=0.500.\n\nIf only Positives and Certain Negatives are included in the Result Set, `score1` can be used as is.\n\nWhen Potential Negatives are included in the Result Set, the number of Negatives varies probabilistically. Assuming all Potential Negatives are included in the Result Set with probability p, the score can be calculated using the probability mass function (pmf) of the binomial distribution as follows:\n\n$$\n\\text{score2}(n, m, p) = \\sum_{k=0}^{m} \\binom{m}{k} p^k (1-p)^{m-k} \\text{score1}(n, k) \n$$\n\nFor example, score2(50,50,1.0) = 0.500, score2(50,50,0.1)=0.910, and score2(50,50,0.01)=0.990.\n\nThe validation score was calculated as the average score2(n, m, 0.2) for each of the 2500 rows in the self-created test.csv equivalent. The total number of patents is approximately 13 million, and the number of Potential Negatives that could be included in the test index is around 75,000, so `p=75000/13000000 ≒ 0.006` seemed appropriate. However, to avoid optimistic validation scores, p=0.2 was used.\n\nThe equivalent of the publication_number in test.csv was prepared as follows:\n- 2500 patents since 1975 and `len(cpc_codes) > 0`\n\nWhen the test index was created by simply filtering patents after 1975, the validation score and LB score seemed slightly off. Analyzing the published train index and LB revealed that design patents starting with US-D were rarely included in the train index and test.csv used for LB calculation. Upon investigating several design patents, I found that cpc_codes were often empty, so I added the filter `len(cpc_codes) > 0`. I also confirmed that patents with empty cpc_codes were not included in the publication_number column of LB's test.csv.\n\nWith the above modifications, the difference between the LB score and the validation score was generally within 0.01.\n\n## Solution\n- Phase 1: Find subquery candidates\n- Phase 2: Generate queries using beam search to maximize `score2(n, m, 0.2)`\n\nTo achieve fast execution speed, as with my own searcher, everything was implemented in C++.\n\n### Phase 1: Finding Subquery Candidates\n- n-shot subquery\n  - A subquery that narrows down to a specific single patent. For example, ti:margaric narrows down to US-2023092870-A1 because it is the only patent containing margaric in the title. Similarly, the cpc_codes A46B13/005 and A47L7/0061 co-occur only in US-2023025335-A1, so `cpc:A46B13/005 AND cpc:A47L7/0061` narrows down to US-2023025335-A1.\n  - Pre-calculated and embedded.\n    - Found that 5.9M patents could be narrowed down to one with a single token, and an additional 6.7M patents could be narrowed down with two tokens.\n- conjunctive subquery\n  - A subquery that includes multiple Positives, no Certain Negatives, and at most `l` Potential Negatives.\n  - Combination of 2-3 tokens using AND. \n    - Explored combinations of cpc_codes, words in title and abstract (frequency <=400,000), claims (frequency <=100,000), and description (frequency <=10,000) using DFS (Depth-First Search).\n    - Explored most cases fully, but some were time-limited.\n  - Avoided Certain Negatives as they significantly reduce the score and facilitate efficient DFS pruning.\n  - Conducted beam search with `l=0` and `l=1` for Phase 2. This improved validation score by about 0.003.\n  - Calculated on-the-fly rather than pre-calculating.\n\n### Phase 2: Beam Search\nThe evaluation metric used was score2(n, m, 0.2), and I performed beam search with a width W for each number of tokens used. For the submission, I set W=100. Increasing W slightly improved the validation score, but even increasing W by a factor of 10 only changed the validation score by about 0.001. Additionally, if subqueries contained common tokens, they were combined to save on the number of tokens used.\n\nFor example, the query for US-7507696-B2 is as follows:  \n`(ti:composition ((detd:coox detd:pearly) OR (clm:dye detd:behentrimoinium))) OR (cpc:A61Q5/12 detd:artichoke detd:genaminox) OR detd:amidoquatsagain`\n\nThis query matches only the 50 target patents of US-7507696-B2 out of approximately 13 million patents. The query is broken down into four subqueries:\n\n- `ti:composition detd:coox detd:pearly`\n- `ti:composition clm:dye detd:behentrimoinium`\n- `cpc:A61Q5/12 detd:artichoke detd:genaminox`\n- `detd:amidoquatsagain`\n\nThe first three subqueries are conjunctive subqueries, and the first and second share the common token ti:composition, so they are combined. The final subquery, detd:amidoquatsagain, is an n-shot subquery and matches only US-8044007-B2.\n\nLike the example query shown, the proportion of perfect queries generated that matched only the 50 target patents out of all patents was about 6%.\n\nThe distribution of score2 in the submission with the best private score is as follows.\n\n![score2_distribution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3097455%2Ff47a9d31197257e42b2dd840194ebaaa%2Fscore2_distribution.png?generation=1721866520258494&alt=media)\n\n## What wasn't used/worked\n- `NOT`\n  - I thought it might be possible to improve the score in the form of `(subquery_a OR ...) NOT (subquery_b OR ...)` by roughly narrowing down the patents and then eliminating the Negatives in the subqueries after NOT, but I didn't have time to implement and test it.\n- Other metaheuristics in Phase 2\n  - I tried replacing candidates using hill climbing or simulated annealing, but simply replacing candidates randomly was inferior to beam search. Moreover, even if I worked hard to improve Phase 2 alone, the validation score seemed to improve by less than 0.01, so I didn't delve too deeply into it.",
    "2935303": "Congratulations on winning the 6th place in this competition. Thanks for sharing your insights.",
    "2936219": "Thanks for an amazing solution with a custom C++ code, and maybe the 1st without the magic!\n\nDoes your custom searcher search the whole patent data parquet in the disk `kaggle/input`, or search only partial negative candidates on RAM? Can I ask what data are on RAM?\n\nCongratulations for the solo gold!",
    "2936890": "Thank you!\n\nThe data after performing similar preprocessing as in the 1st/2nd place solutions is loaded into RAM.\n- The search Index contains the whole patent in patent_metadata.parquet.\n- Convert patents and words into IDs instead of using strings.\n- Create a two-dimensional variable-length list from patents to words and from words to patents.\n  - In the list from patents to words, exclude patents and words that do not appear in the patents included in test.csv.\n  - Similarly, in the list from words to patents, exclude words that do not appear in those patents.\n- And so on.\n\nFor explanatory purposes, I have separated the searcher and solver, but in reality, they are implemented in a single C++ file and are tightly coupled. The preprocessing scripts and the notebook will be organized and made public later, so please refer to those for details.",
    "2937242": "Thanks! I thought accessing the disk multiple times is too slow, but handling the whole patent was also impossible in terms of time and RAM. Amazing! Congratulations again!",
    "2939851": "Nice solution, very similar to mine ! But much more efficient with C++... So you had access to more vocabulary without raising memory errors. Very clean and impressing !\n\nA well deserved cahsprize 🥳",
    "2940228": "Thanks for sharing your solution and congrats on the the prize. \n\nThere is one thing that i don't understand: what is the purpose of distinguish certain negative and potential negative? if it is solely related to how the MAP@50 is calculated, why avoiding certain negative can increase the score? thanks!",
    "2940809": "Thanks!\n\nThere are two reasons:\n- If Certain Negatives are included, the score on the LB decreases. For example, Consider a `query_a` with 20 Positives + 20 Certain Negatives and a `query_b` with 20 Positives + 20 Potential Negatives on the validation index that includes all patents. In this case, the expected value of AP'@50 (≒LB score) for `query_a` is score1(20,20)=0.489. On the other hand, since Potential Negatives are only probabilistically included in the test index, the expected value of AP'@50 for `query_b` will be much higher than 0.489.\n- The calculation efficiency of DFS for conjunctive subqueries in Phase 1 improves. The condition of not including Certain Negatives can be confirmed by checking only Certain Negatives, which are included in about 1% of all patents, making it an effective pruning."
  },
  "source": "meta"
}