{
  "id": 522586,
  "title": "50th Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/filtered-50th-place-solution",
  "author_name": "",
  "post_date": "2024-07-27T01:28:37.703Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm glad that I'm able to hold the last spot in the silver zone at the end and get my first solo silver! Thanks to the USPTO for bringing such a great competition! Participants showed a variety of creative methods in this competition, and it doesn't rely much on GPUs, which is friendly to people who don't have the resources. Once again, hats off to the organizers!</p>\n<h1>Solution Summary</h1>\n<p>Salute to this notebook: <a href=\"https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline\" target=\"_blank\">https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline</a>. Most of my ideas have come from exploring it. Here are some key points in my solution:</p>\n<h2>1. Multiway to Recall Keywords</h2>\n<p>I first found out that increasing the number of keywords would make the query more accurate. So I tried multiple ways to recall keywords then assemble them. This included using TF-IDF on abstract, claims, description and the improved version of <a href=\"https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng\" target=\"_blank\">https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng</a> and <a href=\"https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query\" target=\"_blank\">https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query</a>, but they either use too much time or didn't improve the score. So I ended up just keeping the TF-IDF on cpc and title.</p>\n<h2>2. Genetic Algorithm</h2>\n<p>Due to recalling more than the limited number of keywords, I needed a way to do keyword selection. Simulated annealing doesn't do this very well. (you'll find that removing simulated annealing part from the baseline leads to higher scores, because it is easy to fall into a local solution near the random initial state, this is not as good as having all keywords appear.) So I choose genetic algorithm, which makes the bad keys to be eliminated by natural selection. Repeat this process until the number of keywords is less than the limit. </p>\n<h2>3. Multi-Processing</h2>\n<p>One drawback of the genetic algorithm is that it is time consuming. Noting that whoosh supports concurrency and each line is independent, I used multi-processing to generate the queries, and applied it to process files to save more time. </p>\n<h2>4. Advanced Usage of Whoosh</h2>\n<p>I didn't use the magic of tricking out <code>count_query_tokens</code>, but I did use the advanced tricks of Whoosh, such as using AND to join keywords in different categories, and NOT to filter out patents that have a certain field. Thanks <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104</a>!</p>\n<p>There is still a lot that can be done to improve my method, for example simply increasing the recall number, generation and population size, or combining it with simulated annealing can lead to an improvement in score. But all of these would cause it to exceed the submission time limit (I've been exploring this limit for the last couple of days). It can reach 0.7 - 0.8 in local tests. Here is my code, I'd be happy if you guys explore its full potential!<br>\n<a href=\"https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection\" target=\"_blank\">https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection</a></p>",
  "messages": [
    {
      "id": "2937193",
      "postDate": "07/26/2024 19:32:37",
      "content": "<p>I'm glad that I'm able to hold the last spot in the silver zone at the end and get my first solo silver! Thanks to the USPTO for bringing such a great competition! Participants showed a variety of creative methods in this competition, and it doesn't rely much on GPUs, which is friendly to people who don't have the resources. Once again, hats off to the organizers!</p>\n<h1>Solution Summary</h1>\n<p>Salute to this notebook: <a href=\"https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline\" target=\"_blank\">https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline</a>. Most of my ideas have come from exploring it. Here are some key points in my solution:</p>\n<h2>1. Multiway to Recall Keywords</h2>\n<p>I first found out that increasing the number of keywords would make the query more accurate. So I tried multiple ways to recall keywords then assemble them. This included using TF-IDF on abstract, claims, description and the improved version of <a href=\"https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng\" target=\"_blank\">https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng</a> and <a href=\"https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query\" target=\"_blank\">https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query</a>, but they either use too much time or didn't improve the score. So I ended up just keeping the TF-IDF on cpc and title.</p>\n<h2>2. Genetic Algorithm</h2>\n<p>Due to recalling more than the limited number of keywords, I needed a way to do keyword selection. Simulated annealing doesn't do this very well. (you'll find that removing simulated annealing part from the baseline leads to higher scores, because it is easy to fall into a local solution near the random initial state, this is not as good as having all keywords appear.) So I choose genetic algorithm, which makes the bad keys to be eliminated by natural selection. Repeat this process until the number of keywords is less than the limit. </p>\n<h2>3. Multi-Processing</h2>\n<p>One drawback of the genetic algorithm is that it is time consuming. Noting that whoosh supports concurrency and each line is independent, I used multi-processing to generate the queries, and applied it to process files to save more time. </p>\n<h2>4. Advanced Usage of Whoosh</h2>\n<p>I didn't use the magic of tricking out <code>count_query_tokens</code>, but I did use the advanced tricks of Whoosh, such as using AND to join keywords in different categories, and NOT to filter out patents that have a certain field. Thanks <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104</a>!</p>\n<p>There is still a lot that can be done to improve my method, for example simply increasing the recall number, generation and population size, or combining it with simulated annealing can lead to an improvement in score. But all of these would cause it to exceed the submission time limit (I've been exploring this limit for the last couple of days). It can reach 0.7 - 0.8 in local tests. Here is my code, I'd be happy if you guys explore its full potential!<br>\n<a href=\"https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection\" target=\"_blank\">https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection</a></p>",
      "rawMarkdown": "I'm glad that I'm able to hold the last spot in the silver zone at the end and get my first solo silver! Thanks to the USPTO for bringing such a great competition! Participants showed a variety of creative methods in this competition, and it doesn't rely much on GPUs, which is friendly to people who don't have the resources. Once again, hats off to the organizers!\n\n# Solution Summary\nSalute to this notebook: [https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline](https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline). Most of my ideas have come from exploring it. Here are some key points in my solution:\n\n## 1. Multiway to Recall Keywords\nI first found out that increasing the number of keywords would make the query more accurate. So I tried multiple ways to recall keywords then assemble them. This included using TF-IDF on abstract, claims, description and the improved version of [https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng](https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng) and [https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query](https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query), but they either use too much time or didn't improve the score. So I ended up just keeping the TF-IDF on cpc and title.\n## 2. Genetic Algorithm\nDue to recalling more than the limited number of keywords, I needed a way to do keyword selection. Simulated annealing doesn't do this very well. (you'll find that removing simulated annealing part from the baseline leads to higher scores, because it is easy to fall into a local solution near the random initial state, this is not as good as having all keywords appear.) So I choose genetic algorithm, which makes the bad keys to be eliminated by natural selection. Repeat this process until the number of keywords is less than the limit. \n## 3. Multi-Processing\nOne drawback of the genetic algorithm is that it is time consuming. Noting that whoosh supports concurrency and each line is independent, I used multi-processing to generate the queries, and applied it to process files to save more time. \n## 4. Advanced Usage of Whoosh\nI didn't use the magic of tricking out `count_query_tokens`, but I did use the advanced tricks of Whoosh, such as using AND to join keywords in different categories, and NOT to filter out patents that have a certain field. Thanks [https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104)!\n\nThere is still a lot that can be done to improve my method, for example simply increasing the recall number, generation and population size, or combining it with simulated annealing can lead to an improvement in score. But all of these would cause it to exceed the submission time limit (I've been exploring this limit for the last couple of days). It can reach 0.7 - 0.8 in local tests. Here is my code, I'd be happy if you guys explore its full potential!\n[https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection](https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection)",
      "votes": null
    },
    {
      "id": "2937369",
      "postDate": "07/27/2024 02:47:15",
      "content": "<p>Congratulations on your first solo silver! Your innovative use of genetic algorithms for keyword selection and multi-processing to save time is truly impressive.thank you for sharing insight  <a href=\"https://www.kaggle.com/huanligong\" target=\"_blank\">@huanligong</a> </p>",
      "rawMarkdown": "Congratulations on your first solo silver! Your innovative use of genetic algorithms for keyword selection and multi-processing to save time is truly impressive.thank you for sharing insight  @huanligong",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2937369,
      "author_name": "aaditshukla",
      "author_url": "",
      "post_date": "07/27/2024 02:47:15",
      "content": "<p>Congratulations on your first solo silver! Your innovative use of genetic algorithms for keyword selection and multi-processing to save time is truly impressive.thank you for sharing insight  <a href=\"https://www.kaggle.com/huanligong\" target=\"_blank\">@huanligong</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2937193": "I'm glad that I'm able to hold the last spot in the silver zone at the end and get my first solo silver! Thanks to the USPTO for bringing such a great competition! Participants showed a variety of creative methods in this competition, and it doesn't rely much on GPUs, which is friendly to people who don't have the resources. Once again, hats off to the organizers!\n\n# Solution Summary\nSalute to this notebook: [https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline](https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline). Most of my ideas have come from exploring it. Here are some key points in my solution:\n\n## 1. Multiway to Recall Keywords\nI first found out that increasing the number of keywords would make the query more accurate. So I tried multiple ways to recall keywords then assemble them. This included using TF-IDF on abstract, claims, description and the improved version of [https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng](https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng) and [https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query](https://www.kaggle.com/code/seshurajup/lb-0-11-uspto-single-cpc-query), but they either use too much time or didn't improve the score. So I ended up just keeping the TF-IDF on cpc and title.\n## 2. Genetic Algorithm\nDue to recalling more than the limited number of keywords, I needed a way to do keyword selection. Simulated annealing doesn't do this very well. (you'll find that removing simulated annealing part from the baseline leads to higher scores, because it is easy to fall into a local solution near the random initial state, this is not as good as having all keywords appear.) So I choose genetic algorithm, which makes the bad keys to be eliminated by natural selection. Repeat this process until the number of keywords is less than the limit. \n## 3. Multi-Processing\nOne drawback of the genetic algorithm is that it is time consuming. Noting that whoosh supports concurrency and each line is independent, I used multi-processing to generate the queries, and applied it to process files to save more time. \n## 4. Advanced Usage of Whoosh\nI didn't use the magic of tricking out `count_query_tokens`, but I did use the advanced tricks of Whoosh, such as using AND to join keywords in different categories, and NOT to filter out patents that have a certain field. Thanks [https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/516104)!\n\nThere is still a lot that can be done to improve my method, for example simply increasing the recall number, generation and population size, or combining it with simulated annealing can lead to an improvement in score. But all of these would cause it to exceed the submission time limit (I've been exploring this limit for the last couple of days). It can reach 0.7 - 0.8 in local tests. Here is my code, I'd be happy if you guys explore its full potential!\n[https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection](https://www.kaggle.com/code/huanligong/uspto-genetic-algorithm-for-keyword-selection)",
    "2937369": "Congratulations on your first solo silver! Your innovative use of genetic algorithms for keyword selection and multi-processing to save time is truly impressive.thank you for sharing insight  @huanligong"
  },
  "source": "meta"
}