{
  "id": 522233,
  "title": "1st place solution",
  "url": "/competitions/uspto-explainable-ai/discussion/522233",
  "author_name": "tk",
  "post_date": "2024-07-25T06:01:41.880000",
  "votes": 39,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thanks to the hosts for this interesting competition. Congratulations to the winning teams. Thanks also to <a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> for competing with me.</p>\n<h1>summary</h1>\n<ul>\n<li>Simulated Annealing</li>\n<li>Only AND, OR</li>\n<li>Omission of AND using -</li>\n</ul>\n<h1>query</h1>\n<p>The format of the query is as follows:<br>\n(ti:word1-ti:word2-detd:word3-…-cpc:wordN) OR …</p>\n<ul>\n<li>The query is composed only of AND, OR</li>\n<li>By connecting words with -, the AND token can be omitted</li>\n<li>cpc cannot be omitted with -, so cpc is placed last</li>\n<li>Use all words for cpc, title, abstract</li>\n<li>Delete words with high frequency (100,000 or more) for claim, description</li>\n</ul>\n<h1>candidate generation</h1>\n<p>Define a sequence of words connected by AND as a subquery.<br>\nGenerate candidates for subqueries to be used in the query through the following steps:</p>\n<ol>\n<li>Generate a set of words to be used in subqueries<ul>\n<li>All words possessed by a single target</li>\n<li>Common set of words possessed by two targets</li></ul></li>\n<li>Sort words by the number of elements</li>\n<li>Add words until the patent set consists only of targets, and make it a subquery<ul>\n<li>The patent set is obtained by taking the common set of each word</li>\n<li>If there are non-targets after combining all the words, that subquery is not a candidate</li></ul></li>\n</ol>\n<h2>tips for improvement</h2>\n<ul>\n<li>Reduce computational complexity by adding words in ascending order of elements<ul>\n<li>The complexity of calculating the common set of two sets s, t is min(len(s), len(t))</li></ul></li>\n<li>Speed up the calculation of the common set using cupy<ul>\n<li>2-3 times faster compared to using set(a) &amp; set(b)</li>\n<li>cp.intersect1d(array1, array2)</li>\n<li>On the final day, speeding up with this allowed using all cpc, title, abstract, and the rank improved from 3rd to 1st</li></ul></li>\n<li>Reduce memory usage by placing only patents appearing in test.csv (2500*50) and the words those patents possess in memory</li>\n</ul>\n<h1>Simulated Annealing</h1>\n<ul>\n<li>Combine subqueries with OR</li>\n<li>Neighborhood<ul>\n<li>50% chance to add one unused subquery</li>\n<li>50% chance to remove one used subquery</li></ul></li>\n<li>Score function<ul>\n<li>The number of targets included in the search results of the query</li>\n<li>Consider only the number of targets as candidates are subqueries with zero non-targets</li></ul></li>\n<li>Duplicate removal<ul>\n<li>Reduce the number of candidates by removing duplicates, as subqueries with the same target set do not need multiple candidates</li></ul></li>\n</ul>\n<h1>code</h1>\n<p><a href=\"https://www.kaggle.com/code/tanakar/sa-cpc-title-abst-clm-desc-10-5-allpub-cupy-sub\" target=\"_blank\">https://www.kaggle.com/code/tanakar/sa-cpc-title-abst-clm-desc-10-5-allpub-cupy-sub</a></p>",
  "messages": [
    {
      "id": 2935272,
      "postDate": "2024-07-25T06:01:41.880Z",
      "content": "<p>Thanks to the hosts for this interesting competition. Congratulations to the winning teams. Thanks also to <a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> for competing with me.</p>\n<h1>summary</h1>\n<ul>\n<li>Simulated Annealing</li>\n<li>Only AND, OR</li>\n<li>Omission of AND using -</li>\n</ul>\n<h1>query</h1>\n<p>The format of the query is as follows:<br>\n(ti:word1-ti:word2-detd:word3-…-cpc:wordN) OR …</p>\n<ul>\n<li>The query is composed only of AND, OR</li>\n<li>By connecting words with -, the AND token can be omitted</li>\n<li>cpc cannot be omitted with -, so cpc is placed last</li>\n<li>Use all words for cpc, title, abstract</li>\n<li>Delete words with high frequency (100,000 or more) for claim, description</li>\n</ul>\n<h1>candidate generation</h1>\n<p>Define a sequence of words connected by AND as a subquery.<br>\nGenerate candidates for subqueries to be used in the query through the following steps:</p>\n<ol>\n<li>Generate a set of words to be used in subqueries<ul>\n<li>All words possessed by a single target</li>\n<li>Common set of words possessed by two targets</li></ul></li>\n<li>Sort words by the number of elements</li>\n<li>Add words until the patent set consists only of targets, and make it a subquery<ul>\n<li>The patent set is obtained by taking the common set of each word</li>\n<li>If there are non-targets after combining all the words, that subquery is not a candidate</li></ul></li>\n</ol>\n<h2>tips for improvement</h2>\n<ul>\n<li>Reduce computational complexity by adding words in ascending order of elements<ul>\n<li>The complexity of calculating the common set of two sets s, t is min(len(s), len(t))</li></ul></li>\n<li>Speed up the calculation of the common set using cupy<ul>\n<li>2-3 times faster compared to using set(a) &amp; set(b)</li>\n<li>cp.intersect1d(array1, array2)</li>\n<li>On the final day, speeding up with this allowed using all cpc, title, abstract, and the rank improved from 3rd to 1st</li></ul></li>\n<li>Reduce memory usage by placing only patents appearing in test.csv (2500*50) and the words those patents possess in memory</li>\n</ul>\n<h1>Simulated Annealing</h1>\n<ul>\n<li>Combine subqueries with OR</li>\n<li>Neighborhood<ul>\n<li>50% chance to add one unused subquery</li>\n<li>50% chance to remove one used subquery</li></ul></li>\n<li>Score function<ul>\n<li>The number of targets included in the search results of the query</li>\n<li>Consider only the number of targets as candidates are subqueries with zero non-targets</li></ul></li>\n<li>Duplicate removal<ul>\n<li>Reduce the number of candidates by removing duplicates, as subqueries with the same target set do not need multiple candidates</li></ul></li>\n</ul>\n<h1>code</h1>\n<p><a href=\"https://www.kaggle.com/code/tanakar/sa-cpc-title-abst-clm-desc-10-5-allpub-cupy-sub\" target=\"_blank\">https://www.kaggle.com/code/tanakar/sa-cpc-title-abst-clm-desc-10-5-allpub-cupy-sub</a></p>",
      "rawMarkdown": "\nThanks to the hosts for this interesting competition. Congratulations to the winning teams. Thanks also to [@tomyanabe](https://www.kaggle.com/tomyanabe) for competing with me.\n\n# summary\n\n* Simulated Annealing\n* Only AND, OR\n* Omission of AND using -\n\n# query\n\nThe format of the query is as follows:\n(ti:word1-ti:word2-detd:word3-...-cpc:wordN) OR ...\n\n* The query is composed only of AND, OR\n* By connecting words with -, the AND token can be omitted\n* cpc cannot be omitted with -, so cpc is placed last\n* Use all words for cpc, title, abstract\n* Delete words with high frequency (100,000 or more) for claim, description\n\n# candidate generation\nDefine a sequence of words connected by AND as a subquery.\nGenerate candidates for subqueries to be used in the query through the following steps:\n\n1. Generate a set of words to be used in subqueries\n    * All words possessed by a single target\n    * Common set of words possessed by two targets\n2. Sort words by the number of elements\n3. Add words until the patent set consists only of targets, and make it a subquery\n    * The patent set is obtained by taking the common set of each word\n    * If there are non-targets after combining all the words, that subquery is not a candidate\n\n## tips for improvement\n* Reduce computational complexity by adding words in ascending order of elements\n    * The complexity of calculating the common set of two sets s, t is min(len(s), len(t))\n* Speed up the calculation of the common set using cupy\n    * 2-3 times faster compared to using set(a) & set(b)\n    * cp.intersect1d(array1, array2)\n    * On the final day, speeding up with this allowed using all cpc, title, abstract, and the rank improved from 3rd to 1st\n* Reduce memory usage by placing only patents appearing in test.csv (2500*50) and the words those patents possess in memory\n\n# Simulated Annealing\n* Combine subqueries with OR\n* Neighborhood\n    * 50% chance to add one unused subquery\n    * 50% chance to remove one used subquery\n* Score function\n    * The number of targets included in the search results of the query\n    * Consider only the number of targets as candidates are subqueries with zero non-targets\n* Duplicate removal\n    * Reduce the number of candidates by removing duplicates, as subqueries with the same target set do not need multiple candidates\n\n# code\nhttps://www.kaggle.com/code/tanakar/sa-cpc-title-abst-clm-desc-10-5-allpub-cupy-sub",
      "votes": 39
    },
    {
      "id": 3069131,
      "postDate": "2024-12-11T05:08:07.717Z",
      "content": "<p>It helped me a lot</p>",
      "rawMarkdown": "It helped me a lot"
    },
    {
      "id": 2955634,
      "postDate": "2024-08-11T11:57:52.100Z",
      "content": "<p>Amazing Work, will definitely want you to be my mentor.</p>",
      "rawMarkdown": "Amazing Work, will definitely want you to be my mentor."
    },
    {
      "id": 2946329,
      "postDate": "2024-08-04T09:34:59.327Z",
      "content": "<p>NICE, I study English and Data Science</p>",
      "rawMarkdown": "NICE, I study English and Data Science"
    },
    {
      "id": 2939568,
      "postDate": "2024-07-29T09:52:05.130Z",
      "content": "<p>masterpiece.</p>",
      "rawMarkdown": "masterpiece."
    },
    {
      "id": 2939128,
      "postDate": "2024-07-28T18:59:25.950Z",
      "content": "<p>good job broo</p>",
      "rawMarkdown": "good job broo"
    },
    {
      "id": 2935299,
      "postDate": "2024-07-25T06:28:02.660Z",
      "content": "<p>Congratulations on winning the top place in this competition. Thanks for sharing the details of your solution. Looking forward to reviewing your code. </p>",
      "rawMarkdown": "Congratulations on winning the top place in this competition. Thanks for sharing the details of your solution. Looking forward to reviewing your code. "
    },
    {
      "id": 2940836,
      "postDate": "2024-07-30T13:52:11.993Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 3042920,
      "postDate": "2024-11-11T22:50:27.380Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 2975283,
      "postDate": "2024-08-31T17:18:23.153Z",
      "content": "<p>thank you for sharing </p>",
      "rawMarkdown": "thank you for sharing "
    },
    {
      "id": 2935294,
      "postDate": "2024-07-25T06:24:50.007Z",
      "content": "<p>thank you for sharing <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> </p>",
      "rawMarkdown": "thank you for sharing @tanakar "
    }
  ],
  "comments": [
    {
      "id": 3069131,
      "author_name": "Labmus Qooraf",
      "author_url": "",
      "post_date": "2024-12-11T05:08:07.717000",
      "content": "<p>It helped me a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2955634,
      "author_name": "Ateeq Ur Rehman",
      "author_url": "",
      "post_date": "2024-08-11T11:57:52.100000",
      "content": "<p>Amazing Work, will definitely want you to be my mentor.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2946329,
      "author_name": "Suzuki Shunta",
      "author_url": "",
      "post_date": "2024-08-04T09:34:59.327000",
      "content": "<p>NICE, I study English and Data Science</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2939568,
      "author_name": "Qiao Shang",
      "author_url": "",
      "post_date": "2024-07-29T09:52:05.130000",
      "content": "<p>masterpiece.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2939128,
      "author_name": "Abdullah Ali Abdullah",
      "author_url": "",
      "post_date": "2024-07-28T18:59:25.950000",
      "content": "<p>good job broo</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935299,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-07-25T06:28:02.660000",
      "content": "<p>Congratulations on winning the top place in this competition. Thanks for sharing the details of your solution. Looking forward to reviewing your code. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940836,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T13:52:11.993000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3042920,
      "author_name": "Iris W",
      "author_url": "",
      "post_date": "2024-11-11T22:50:27.380000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2975283,
      "author_name": "ginger",
      "author_url": "",
      "post_date": "2024-08-31T17:18:23.153000",
      "content": "<p>thank you for sharing </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935294,
      "author_name": "Aadit Shukla",
      "author_url": "",
      "post_date": "2024-07-25T06:24:50.007000",
      "content": "<p>thank you for sharing <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935272": "\nThanks to the hosts for this interesting competition. Congratulations to the winning teams. Thanks also to [@tomyanabe](https://www.kaggle.com/tomyanabe) for competing with me.\n\n# summary\n\n* Simulated Annealing\n* Only AND, OR\n* Omission of AND using -\n\n# query\n\nThe format of the query is as follows:\n(ti:word1-ti:word2-detd:word3-...-cpc:wordN) OR ...\n\n* The query is composed only of AND, OR\n* By connecting words with -, the AND token can be omitted\n* cpc cannot be omitted with -, so cpc is placed last\n* Use all words for cpc, title, abstract\n* Delete words with high frequency (100,000 or more) for claim, description\n\n# candidate generation\nDefine a sequence of words connected by AND as a subquery.\nGenerate candidates for subqueries to be used in the query through the following steps:\n\n1. Generate a set of words to be used in subqueries\n    * All words possessed by a single target\n    * Common set of words possessed by two targets\n2. Sort words by the number of elements\n3. Add words until the patent set consists only of targets, and make it a subquery\n    * The patent set is obtained by taking the common set of each word\n    * If there are non-targets after combining all the words, that subquery is not a candidate\n\n## tips for improvement\n* Reduce computational complexity by adding words in ascending order of elements\n    * The complexity of calculating the common set of two sets s, t is min(len(s), len(t))\n* Speed up the calculation of the common set using cupy\n    * 2-3 times faster compared to using set(a) & set(b)\n    * cp.intersect1d(array1, array2)\n    * On the final day, speeding up with this allowed using all cpc, title, abstract, and the rank improved from 3rd to 1st\n* Reduce memory usage by placing only patents appearing in test.csv (2500*50) and the words those patents possess in memory\n\n# Simulated Annealing\n* Combine subqueries with OR\n* Neighborhood\n    * 50% chance to add one unused subquery\n    * 50% chance to remove one used subquery\n* Score function\n    * The number of targets included in the search results of the query\n    * Consider only the number of targets as candidates are subqueries with zero non-targets\n* Duplicate removal\n    * Reduce the number of candidates by removing duplicates, as subqueries with the same target set do not need multiple candidates\n\n# code\nhttps://www.kaggle.com/code/tanakar/sa-cpc-title-abst-clm-desc-10-5-allpub-cupy-sub",
    "3069131": "It helped me a lot",
    "2955634": "Amazing Work, will definitely want you to be my mentor.",
    "2946329": "NICE, I study English and Data Science",
    "2939568": "masterpiece.",
    "2939128": "good job broo",
    "2935299": "Congratulations on winning the top place in this competition. Thanks for sharing the details of your solution. Looking forward to reviewing your code. ",
    "2940836": "",
    "3042920": "Thanks for sharing!",
    "2975283": "thank you for sharing ",
    "2935294": "thank you for sharing @tanakar "
  }
}