{
  "id": 522257,
  "title": "8th place solution",
  "url": "/competitions/uspto-explainable-ai/discussion/522257",
  "author_name": "Oleg Kokorin",
  "post_date": "2024-07-25T08:21:29.271000",
  "votes": 14,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulations to the winning teams and thanks to the organizers for hosting this competition!</p>\n<p><strong>Solution summary:</strong><br>\nI create a bag of search tokens for each of the 50 neighboring patents. For that, I take all word tokens from CPC codes, titles, abstracts, claims and descriptions For CPC codes and title words I also take all possible cross-combinations. I take all 2-word combinations from titles, and all 2-code combinations of CPC codes. Then, for abstracts, claims, and descriptions I take 2-grams and 3-grams.<br>\nFor every search token I have a global counter that shows how many patents could possibly be found by the search token. <br>\nI use a greedy algorithm to build a query. A bag of the 50 patent ids is created. I count how many patents in the bag would be found by every search token from every patent in the bag. I use the global counters to reject any token that could add more than 1 False Positive. Once a search token is added to the query, the patents that are going to be found by the subquery are removed from the bag.<br>\n<strong>The Magic:</strong><br>\nIt's not really a magic. To find 2-grams and 3-grams I use subqueries like this:</p>\n<pre><code>\n\n</code></pre>\n<p>The question mark matches a space (as it should), while both of these subqueries count as a single token (as they should).<br>\nThis approach is quite consuming in terms of the memory and compute required. A lot of time was spent to optimize all parts of this pipeline, starting from creating the global counters and to the submission notebook. I was not able to use all possible 2- and 3-grams for descriptions. So when creating these n-grams, I would reject an n-gram if any of its parts is found in more than 1,500,000 patents.<br>\n<strong>Validation:</strong><br>\nI created my own index with 10K anchor patents, and eventually created another 8K index, which gave me good correlation. The difference between LV and LB was 0.03-0.04 for the 0.6-0.7 scores, 0.02 for 0.8 scores and about 0.01 for 0.9+ scores.<br>\n<strong>What didn't work:</strong><br>\nI spent quite a lot of time trying to build effective queries for design patents (they are extremely similar and have no CPC codes), only to learn later than the test set didn't contain any.<br>\nI also spent time to learn how whoosh orders the query results trying to make it order the results in a better way. A single term query like 'clm:mobile' is scored via TF-IDF, as expected. A query with a wildcard like 'clm:\"mobile?device\"' is scored in an arcane way, definitely not TF-IDF.<br>\nI also tried more sophisticated approaches to building the queries (as opposed to the greedy algorithm). That became irrelevant when I created those huge global counters.</p>",
  "messages": [
    {
      "id": 2935413,
      "postDate": "2024-07-25T08:21:29.270Z",
      "content": "<p>Congratulations to the winning teams and thanks to the organizers for hosting this competition!</p>\n<p><strong>Solution summary:</strong><br>\nI create a bag of search tokens for each of the 50 neighboring patents. For that, I take all word tokens from CPC codes, titles, abstracts, claims and descriptions For CPC codes and title words I also take all possible cross-combinations. I take all 2-word combinations from titles, and all 2-code combinations of CPC codes. Then, for abstracts, claims, and descriptions I take 2-grams and 3-grams.<br>\nFor every search token I have a global counter that shows how many patents could possibly be found by the search token. <br>\nI use a greedy algorithm to build a query. A bag of the 50 patent ids is created. I count how many patents in the bag would be found by every search token from every patent in the bag. I use the global counters to reject any token that could add more than 1 False Positive. Once a search token is added to the query, the patents that are going to be found by the subquery are removed from the bag.<br>\n<strong>The Magic:</strong><br>\nIt's not really a magic. To find 2-grams and 3-grams I use subqueries like this:</p>\n<pre><code>\n\n</code></pre>\n<p>The question mark matches a space (as it should), while both of these subqueries count as a single token (as they should).<br>\nThis approach is quite consuming in terms of the memory and compute required. A lot of time was spent to optimize all parts of this pipeline, starting from creating the global counters and to the submission notebook. I was not able to use all possible 2- and 3-grams for descriptions. So when creating these n-grams, I would reject an n-gram if any of its parts is found in more than 1,500,000 patents.<br>\n<strong>Validation:</strong><br>\nI created my own index with 10K anchor patents, and eventually created another 8K index, which gave me good correlation. The difference between LV and LB was 0.03-0.04 for the 0.6-0.7 scores, 0.02 for 0.8 scores and about 0.01 for 0.9+ scores.<br>\n<strong>What didn't work:</strong><br>\nI spent quite a lot of time trying to build effective queries for design patents (they are extremely similar and have no CPC codes), only to learn later than the test set didn't contain any.<br>\nI also spent time to learn how whoosh orders the query results trying to make it order the results in a better way. A single term query like 'clm:mobile' is scored via TF-IDF, as expected. A query with a wildcard like 'clm:\"mobile?device\"' is scored in an arcane way, definitely not TF-IDF.<br>\nI also tried more sophisticated approaches to building the queries (as opposed to the greedy algorithm). That became irrelevant when I created those huge global counters.</p>",
      "rawMarkdown": "Congratulations to the winning teams and thanks to the organizers for hosting this competition!\n\n**Solution summary:**\nI create a bag of search tokens for each of the 50 neighboring patents. For that, I take all word tokens from CPC codes, titles, abstracts, claims and descriptions For CPC codes and title words I also take all possible cross-combinations. I take all 2-word combinations from titles, and all 2-code combinations of CPC codes. Then, for abstracts, claims, and descriptions I take 2-grams and 3-grams.\nFor every search token I have a global counter that shows how many patents could possibly be found by the search token. \nI use a greedy algorithm to build a query. A bag of the 50 patent ids is created. I count how many patents in the bag would be found by every search token from every patent in the bag. I use the global counters to reject any token that could add more than 1 False Positive. Once a search token is added to the query, the patents that are going to be found by the subquery are removed from the bag.\n**The Magic:**\nIt's not really a magic. To find 2-grams and 3-grams I use subqueries like this:\n```python\n'clm:\"mobile?device\"'\n'detd:\"alice?met?bob\"'\n\n```The question mark matches a space (as it should), while both of these subqueries count as a single token (as they should).\nThis approach is quite consuming in terms of the memory and compute required. A lot of time was spent to optimize all parts of this pipeline, starting from creating the global counters and to the submission notebook. I was not able to use all possible 2- and 3-grams for descriptions. So when creating these n-grams, I would reject an n-gram if any of its parts is found in more than 1,500,000 patents.\n**Validation:**\nI created my own index with 10K anchor patents, and eventually created another 8K index, which gave me good correlation. The difference between LV and LB was 0.03-0.04 for the 0.6-0.7 scores, 0.02 for 0.8 scores and about 0.01 for 0.9+ scores.\n**What didn't work:**\nI spent quite a lot of time trying to build effective queries for design patents (they are extremely similar and have no CPC codes), only to learn later than the test set didn't contain any.\nI also spent time to learn how whoosh orders the query results trying to make it order the results in a better way. A single term query like 'clm:mobile' is scored via TF-IDF, as expected. A query with a wildcard like 'clm:\"mobile?device\"' is scored in an arcane way, definitely not TF-IDF.\nI also tried more sophisticated approaches to building the queries (as opposed to the greedy algorithm). That became irrelevant when I created those huge global counters.\n\n",
      "votes": 14
    },
    {
      "id": 2936514,
      "postDate": "2024-07-26T08:14:37.280Z",
      "content": "<p>Great catch! Congratulations!</p>",
      "rawMarkdown": "Great catch! Congratulations!"
    },
    {
      "id": 2935480,
      "postDate": "2024-07-25T10:15:24.333Z",
      "content": "<p>Good job! Amazing!</p>",
      "rawMarkdown": "Good job! Amazing!"
    },
    {
      "id": 2940835,
      "postDate": "2024-07-30T13:50:34.750Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2936514,
      "author_name": "Nick Alymov",
      "author_url": "",
      "post_date": "2024-07-26T08:14:37.280000",
      "content": "<p>Great catch! Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935480,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-07-25T10:15:24.333000",
      "content": "<p>Good job! Amazing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940835,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T13:50:34.750000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935413": "Congratulations to the winning teams and thanks to the organizers for hosting this competition!\n\n**Solution summary:**\nI create a bag of search tokens for each of the 50 neighboring patents. For that, I take all word tokens from CPC codes, titles, abstracts, claims and descriptions For CPC codes and title words I also take all possible cross-combinations. I take all 2-word combinations from titles, and all 2-code combinations of CPC codes. Then, for abstracts, claims, and descriptions I take 2-grams and 3-grams.\nFor every search token I have a global counter that shows how many patents could possibly be found by the search token. \nI use a greedy algorithm to build a query. A bag of the 50 patent ids is created. I count how many patents in the bag would be found by every search token from every patent in the bag. I use the global counters to reject any token that could add more than 1 False Positive. Once a search token is added to the query, the patents that are going to be found by the subquery are removed from the bag.\n**The Magic:**\nIt's not really a magic. To find 2-grams and 3-grams I use subqueries like this:\n```python\n'clm:\"mobile?device\"'\n'detd:\"alice?met?bob\"'\n\n```The question mark matches a space (as it should), while both of these subqueries count as a single token (as they should).\nThis approach is quite consuming in terms of the memory and compute required. A lot of time was spent to optimize all parts of this pipeline, starting from creating the global counters and to the submission notebook. I was not able to use all possible 2- and 3-grams for descriptions. So when creating these n-grams, I would reject an n-gram if any of its parts is found in more than 1,500,000 patents.\n**Validation:**\nI created my own index with 10K anchor patents, and eventually created another 8K index, which gave me good correlation. The difference between LV and LB was 0.03-0.04 for the 0.6-0.7 scores, 0.02 for 0.8 scores and about 0.01 for 0.9+ scores.\n**What didn't work:**\nI spent quite a lot of time trying to build effective queries for design patents (they are extremely similar and have no CPC codes), only to learn later than the test set didn't contain any.\nI also spent time to learn how whoosh orders the query results trying to make it order the results in a better way. A single term query like 'clm:mobile' is scored via TF-IDF, as expected. A query with a wildcard like 'clm:\"mobile?device\"' is scored in an arcane way, definitely not TF-IDF.\nI also tried more sophisticated approaches to building the queries (as opposed to the greedy algorithm). That became irrelevant when I created those huge global counters.\n\n",
    "2936514": "Great catch! Congratulations!",
    "2935480": "Good job! Amazing!",
    "2940835": ""
  }
}