{
  "id": 522639,
  "title": "3rd Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/c-number-3rd-place-solution",
  "author_name": "",
  "post_date": "2024-08-07T06:09:09.353Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I would like to express my sincere gratitude to the organizers for creating this fascinating optimization challenge, and to the dedicated competitors who worked tirelessly to fix Whoosh, ensuring a valid competition.</p>\n<h1>Query Optimization</h1>\n<p>As many participants have pointed out, tokens can be AND-concatenated using the following method without increasing the \"number of tokens\" as measured by whoosh_utils.count_query_tokens:<br>\n<code>ti:token1-token2</code></p>\n<p>The final query consists of these subqueries OR-concatenated, such as:<br>\n<code>(ti:token1-token2) OR (ab:token3-token4-token5) OR …</code></p>\n<p>Regrettably, I was unable to discover the stronger magic that incorporate tokens from different fields into a single subquery.</p>\n<h1>Subquery Search</h1>\n<p>For each sample, up to several thousand subquery candidates were generated. For every small subset of patents (typically size(subset) ≤ 3), tokens common to all patents in the subset were selected and adopted.</p>\n<h1>Subquery Selection</h1>\n<p>Mixed Integer Programming (MIP) solvers were employed to determine the optimal query. In this context, \"optimal\" refers to maximizing the number of target patents found using a random test_index that covers the same number of patents as the train_index. More details can be found in the shared notebook.</p>\n<h1>Key Strategies</h1>\n<ol>\n<li>Implementation of subquery search in C++ with multithreading and aggressive algorithm optimization for improved search speed.</li>\n<li>Consideration of all tokens found in Title, Claim, and CPC fields, while using only the 100,000 least frequent tokens for Abstract and Description fields to reduce computational complexity.</li>\n</ol>\n<h1>Areas for Improvement</h1>\n<ol>\n<li>Utilizing the stronger magic.</li>\n<li>Implementation with CUDA for more intensive search.</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/code/cnumber/uspto-bestsub\" target=\"_blank\">best submission with neater code</a><br>\n<a href=\"https://www.kaggle.com/datasets/cnumber/uspto-preprocess\" target=\"_blank\">codes for preprocessing</a></p>",
  "messages": [
    {
      "id": "2937534",
      "postDate": "07/27/2024 07:22:41",
      "content": "<p>I would like to express my sincere gratitude to the organizers for creating this fascinating optimization challenge, and to the dedicated competitors who worked tirelessly to fix Whoosh, ensuring a valid competition.</p>\n<h1>Query Optimization</h1>\n<p>As many participants have pointed out, tokens can be AND-concatenated using the following method without increasing the \"number of tokens\" as measured by whoosh_utils.count_query_tokens:<br>\n<code>ti:token1-token2</code></p>\n<p>The final query consists of these subqueries OR-concatenated, such as:<br>\n<code>(ti:token1-token2) OR (ab:token3-token4-token5) OR …</code></p>\n<p>Regrettably, I was unable to discover the stronger magic that incorporate tokens from different fields into a single subquery.</p>\n<h1>Subquery Search</h1>\n<p>For each sample, up to several thousand subquery candidates were generated. For every small subset of patents (typically size(subset) ≤ 3), tokens common to all patents in the subset were selected and adopted.</p>\n<h1>Subquery Selection</h1>\n<p>Mixed Integer Programming (MIP) solvers were employed to determine the optimal query. In this context, \"optimal\" refers to maximizing the number of target patents found using a random test_index that covers the same number of patents as the train_index. More details can be found in the shared notebook.</p>\n<h1>Key Strategies</h1>\n<ol>\n<li>Implementation of subquery search in C++ with multithreading and aggressive algorithm optimization for improved search speed.</li>\n<li>Consideration of all tokens found in Title, Claim, and CPC fields, while using only the 100,000 least frequent tokens for Abstract and Description fields to reduce computational complexity.</li>\n</ol>\n<h1>Areas for Improvement</h1>\n<ol>\n<li>Utilizing the stronger magic.</li>\n<li>Implementation with CUDA for more intensive search.</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/code/cnumber/uspto-bestsub\" target=\"_blank\">best submission with neater code</a><br>\n<a href=\"https://www.kaggle.com/datasets/cnumber/uspto-preprocess\" target=\"_blank\">codes for preprocessing</a></p>",
      "rawMarkdown": "I would like to express my sincere gratitude to the organizers for creating this fascinating optimization challenge, and to the dedicated competitors who worked tirelessly to fix Whoosh, ensuring a valid competition.\n\n# Query Optimization\nAs many participants have pointed out, tokens can be AND-concatenated using the following method without increasing the \"number of tokens\" as measured by whoosh_utils.count_query_tokens:\n`ti:token1-token2`\n\nThe final query consists of these subqueries OR-concatenated, such as:\n`(ti:token1-token2) OR (ab:token3-token4-token5) OR …`\n\nRegrettably, I was unable to discover the stronger magic that incorporate tokens from different fields into a single subquery.\n\n# Subquery Search\nFor each sample, up to several thousand subquery candidates were generated. For every small subset of patents (typically size(subset) ≤ 3), tokens common to all patents in the subset were selected and adopted.\n\n# Subquery Selection\nMixed Integer Programming (MIP) solvers were employed to determine the optimal query. In this context, \"optimal\" refers to maximizing the number of target patents found using a random test_index that covers the same number of patents as the train_index. More details can be found in the shared notebook.\n\n# Key Strategies\n1. Implementation of subquery search in C++ with multithreading and aggressive algorithm optimization for improved search speed.\n2. Consideration of all tokens found in Title, Claim, and CPC fields, while using only the 100,000 least frequent tokens for Abstract and Description fields to reduce computational complexity.\n\n# Areas for Improvement\n1. Utilizing the stronger magic.\n2. Implementation with CUDA for more intensive search.\n\n[best submission with neater code](https://www.kaggle.com/code/cnumber/uspto-bestsub)\n[codes for preprocessing](https://www.kaggle.com/datasets/cnumber/uspto-preprocess)",
      "votes": null
    },
    {
      "id": "2938667",
      "postDate": "07/28/2024 11:41:52",
      "content": "<p>thank you for this article ! i wish to achieve a similar rank soon enough !</p>",
      "rawMarkdown": "thank you for this article ! i wish to achieve a similar rank soon enough !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2938667,
      "author_name": "aadiar",
      "author_url": "",
      "post_date": "07/28/2024 11:41:52",
      "content": "<p>thank you for this article ! i wish to achieve a similar rank soon enough !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2937534": "I would like to express my sincere gratitude to the organizers for creating this fascinating optimization challenge, and to the dedicated competitors who worked tirelessly to fix Whoosh, ensuring a valid competition.\n\n# Query Optimization\nAs many participants have pointed out, tokens can be AND-concatenated using the following method without increasing the \"number of tokens\" as measured by whoosh_utils.count_query_tokens:\n`ti:token1-token2`\n\nThe final query consists of these subqueries OR-concatenated, such as:\n`(ti:token1-token2) OR (ab:token3-token4-token5) OR …`\n\nRegrettably, I was unable to discover the stronger magic that incorporate tokens from different fields into a single subquery.\n\n# Subquery Search\nFor each sample, up to several thousand subquery candidates were generated. For every small subset of patents (typically size(subset) ≤ 3), tokens common to all patents in the subset were selected and adopted.\n\n# Subquery Selection\nMixed Integer Programming (MIP) solvers were employed to determine the optimal query. In this context, \"optimal\" refers to maximizing the number of target patents found using a random test_index that covers the same number of patents as the train_index. More details can be found in the shared notebook.\n\n# Key Strategies\n1. Implementation of subquery search in C++ with multithreading and aggressive algorithm optimization for improved search speed.\n2. Consideration of all tokens found in Title, Claim, and CPC fields, while using only the 100,000 least frequent tokens for Abstract and Description fields to reduce computational complexity.\n\n# Areas for Improvement\n1. Utilizing the stronger magic.\n2. Implementation with CUDA for more intensive search.\n\n[best submission with neater code](https://www.kaggle.com/code/cnumber/uspto-bestsub)\n[codes for preprocessing](https://www.kaggle.com/datasets/cnumber/uspto-preprocess)",
    "2938667": "thank you for this article ! i wish to achieve a similar rank soon enough !"
  },
  "source": "meta"
}