{
  "id": 522415,
  "title": "28th Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/cf-28th-place-solution",
  "author_name": "",
  "post_date": "2024-07-26T05:07:22.897Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you to USPTO for organising and hosting this competition, and a big congratulations to all the winning teams.</p>\n<h2>Solution Overview</h2>\n<p>The overall approach was to construct a query in the form: “rare_cpcs_docs OR ((sub-query) AND (sub-query) AND …)”. We made use of a global frequency index for CPCs and an IDF lookup table for Abstract &amp; Claims, which is discussed in detail below.</p>\n<ul>\n<li>We discovered early on that just using CPC codes was very powerful for retrieval. To reduce our search space, we constructed a ‘rare_cpcs_docs’ query, which is an OR-wise concatenation of CPCs. This was only for documents that had very few global hits on CPCs (&lt; 50). Testing locally, this was always around 5-10 documents.</li>\n<li>The aim for each ‘subquery’ is to obtain common concepts to cover the remaining ~40 documents. We did this through a combination of using CPC and keywords extracted from the Abstract &amp; Claims using KeyBERT.</li>\n</ul>\n<p>We focused on rare CPC &amp; rare keywords that shared mostly within 50 docs (i.e CPC with high local hit and low global count, keywords with high local hit and high global IDF), then we greedily picked the best performing entity until all documents were covered (Greedy Set Cover).</p>\n<ul>\n<li>This process was repeated to construct as many iterations of a ‘sub-query’ as possible until we ran over the 50 token limit (no ‘magic’ token limit)</li>\n</ul>\n<p><strong>Generating ‘sub-query’</strong></p>\n<p>We performed a lot of experimentation in generating these sub-queries, but none led to huge gains:</p>\n<ul>\n<li>Using the Title field</li>\n<li>Filtering out overlapping words across sub-queries, e.g., if “electrode” is used in sub-query-1, filter out “negative electrode” so it’s not used in subsequent queries</li>\n<li>ADJ/NEAR within ngram (&gt;=3gram) gives more choices of rare keywords that have higher hit within 50 documents. However, it costs more query term than other operators.</li>\n</ul>\n<p><strong>Constructing Global Indexes</strong></p>\n<p>We analysed the entire set of parquet files (10+ million individual patents) to construct these indexes:</p>\n<ul>\n<li>Frequency count of CPC codes</li>\n<li>Unigram &amp; bigram IDF values for Abstract</li>\n<li>Unigram &amp; bigram IDF values for Claims</li>\n</ul>\n<p>The aim was to use these lookup tables to penalise high frequency terms, thus reducing the number of extraneous documents we retrieved.</p>\n<h3>Other Ideas</h3>\n<ul>\n<li>Our first entry consisted entirely of CPC codes in the form of “cpc:(a OR b OR c OR …)”. The codes were chosen based on the factor of local frequency, i.e., we prioritised codes that covered most of the 50 documents first; a simple greedy set cover algorithm. This simple query scored 0.37 on the leaderboard.</li>\n<li>We also tried to focus the ‘subquery’ on each part of a patent such as (ti:subquery) (clm:subquery)(cpc:subquery). This would help to detect the common patent landscape of each part over 50 docs. Due to the small query size on each subquery, this solution didn’t produce the expected improvements.</li>\n</ul>\n<h3>Tips &amp; Tricks</h3>\n<ul>\n<li>We used Polars + multi-threading to very quickly process and store the 125,000 relevant patents in memory for instant retrieval</li>\n</ul>",
  "messages": [
    {
      "id": "2936366",
      "postDate": "07/26/2024 05:03:22",
      "content": "<p>Thank you to USPTO for organising and hosting this competition, and a big congratulations to all the winning teams.</p>\n<h2>Solution Overview</h2>\n<p>The overall approach was to construct a query in the form: “rare_cpcs_docs OR ((sub-query) AND (sub-query) AND …)”. We made use of a global frequency index for CPCs and an IDF lookup table for Abstract &amp; Claims, which is discussed in detail below.</p>\n<ul>\n<li>We discovered early on that just using CPC codes was very powerful for retrieval. To reduce our search space, we constructed a ‘rare_cpcs_docs’ query, which is an OR-wise concatenation of CPCs. This was only for documents that had very few global hits on CPCs (&lt; 50). Testing locally, this was always around 5-10 documents.</li>\n<li>The aim for each ‘subquery’ is to obtain common concepts to cover the remaining ~40 documents. We did this through a combination of using CPC and keywords extracted from the Abstract &amp; Claims using KeyBERT.</li>\n</ul>\n<p>We focused on rare CPC &amp; rare keywords that shared mostly within 50 docs (i.e CPC with high local hit and low global count, keywords with high local hit and high global IDF), then we greedily picked the best performing entity until all documents were covered (Greedy Set Cover).</p>\n<ul>\n<li>This process was repeated to construct as many iterations of a ‘sub-query’ as possible until we ran over the 50 token limit (no ‘magic’ token limit)</li>\n</ul>\n<p><strong>Generating ‘sub-query’</strong></p>\n<p>We performed a lot of experimentation in generating these sub-queries, but none led to huge gains:</p>\n<ul>\n<li>Using the Title field</li>\n<li>Filtering out overlapping words across sub-queries, e.g., if “electrode” is used in sub-query-1, filter out “negative electrode” so it’s not used in subsequent queries</li>\n<li>ADJ/NEAR within ngram (&gt;=3gram) gives more choices of rare keywords that have higher hit within 50 documents. However, it costs more query term than other operators.</li>\n</ul>\n<p><strong>Constructing Global Indexes</strong></p>\n<p>We analysed the entire set of parquet files (10+ million individual patents) to construct these indexes:</p>\n<ul>\n<li>Frequency count of CPC codes</li>\n<li>Unigram &amp; bigram IDF values for Abstract</li>\n<li>Unigram &amp; bigram IDF values for Claims</li>\n</ul>\n<p>The aim was to use these lookup tables to penalise high frequency terms, thus reducing the number of extraneous documents we retrieved.</p>\n<h3>Other Ideas</h3>\n<ul>\n<li>Our first entry consisted entirely of CPC codes in the form of “cpc:(a OR b OR c OR …)”. The codes were chosen based on the factor of local frequency, i.e., we prioritised codes that covered most of the 50 documents first; a simple greedy set cover algorithm. This simple query scored 0.37 on the leaderboard.</li>\n<li>We also tried to focus the ‘subquery’ on each part of a patent such as (ti:subquery) (clm:subquery)(cpc:subquery). This would help to detect the common patent landscape of each part over 50 docs. Due to the small query size on each subquery, this solution didn’t produce the expected improvements.</li>\n</ul>\n<h3>Tips &amp; Tricks</h3>\n<ul>\n<li>We used Polars + multi-threading to very quickly process and store the 125,000 relevant patents in memory for instant retrieval</li>\n</ul>",
      "rawMarkdown": "Thank you to USPTO for organising and hosting this competition, and a big congratulations to all the winning teams.\n\n## Solution Overview\n\nThe overall approach was to construct a query in the form: “rare_cpcs_docs OR ((sub-query) AND (sub-query) AND …)”. We made use of a global frequency index for CPCs and an IDF lookup table for Abstract & Claims, which is discussed in detail below.\n\n- We discovered early on that just using CPC codes was very powerful for retrieval. To reduce our search space, we constructed a ‘rare_cpcs_docs’ query, which is an OR-wise concatenation of CPCs. This was only for documents that had very few global hits on CPCs (< 50). Testing locally, this was always around 5-10 documents.\n- The aim for each ‘subquery’ is to obtain common concepts to cover the remaining ~40 documents. We did this through a combination of using CPC and keywords extracted from the Abstract & Claims using KeyBERT.\n\nWe focused on rare CPC & rare keywords that shared mostly within 50 docs (i.e CPC with high local hit and low global count, keywords with high local hit and high global IDF), then we greedily picked the best performing entity until all documents were covered (Greedy Set Cover).\n\n- This process was repeated to construct as many iterations of a ‘sub-query’ as possible until we ran over the 50 token limit (no ‘magic’ token limit)\n\n**Generating ‘sub-query’**\n\nWe performed a lot of experimentation in generating these sub-queries, but none led to huge gains:\n\n- Using the Title field\n- Filtering out overlapping words across sub-queries, e.g., if “electrode” is used in sub-query-1, filter out “negative electrode” so it’s not used in subsequent queries\n- ADJ/NEAR within ngram (>=3gram) gives more choices of rare keywords that have higher hit within 50 documents. However, it costs more query term than other operators.\n\n**Constructing Global Indexes**\n\nWe analysed the entire set of parquet files (10+ million individual patents) to construct these indexes:\n\n- Frequency count of CPC codes\n- Unigram & bigram IDF values for Abstract\n- Unigram & bigram IDF values for Claims\n\nThe aim was to use these lookup tables to penalise high frequency terms, thus reducing the number of extraneous documents we retrieved.\n\n### Other Ideas\n\n- Our first entry consisted entirely of CPC codes in the form of “cpc:(a OR b OR c OR …)”. The codes were chosen based on the factor of local frequency, i.e., we prioritised codes that covered most of the 50 documents first; a simple greedy set cover algorithm. This simple query scored 0.37 on the leaderboard.\n- We also tried to focus the ‘subquery’ on each part of a patent such as (ti:subquery) (clm:subquery)(cpc:subquery). This would help to detect the common patent landscape of each part over 50 docs. Due to the small query size on each subquery, this solution didn’t produce the expected improvements.\n\n### Tips & Tricks\n\n- We used Polars + multi-threading to very quickly process and store the 125,000 relevant patents in memory for instant retrieval",
      "votes": null
    },
    {
      "id": "2936369",
      "postDate": "07/26/2024 05:05:55",
      "content": "<p>Thank you for sharing your detailed solution and innovative approach—it's insightful how you utilized rare CPC codes and keywords to refine your search space. Congratulations on your impressive performance and contributions to the competition! <a href=\"https://www.kaggle.com/gagankaur16\" target=\"_blank\">@gagankaur16</a> </p>",
      "rawMarkdown": "Thank you for sharing your detailed solution and innovative approach—it's insightful how you utilized rare CPC codes and keywords to refine your search space. Congratulations on your impressive performance and contributions to the competition! @gagankaur16",
      "votes": null
    },
    {
      "id": "2936398",
      "postDate": "07/26/2024 06:09:48",
      "content": "<p><a href=\"https://www.kaggle.com/gagankaur16\" target=\"_blank\">@gagankaur16</a> Thanks for sharing your approach, and congrats on your medal! Can you explain how you did instant retrieval with Polars?</p>",
      "rawMarkdown": "gagankaur16 Thanks for sharing your approach, and congrats on your medal! Can you explain how you did instant retrieval with Polars?",
      "votes": null
    },
    {
      "id": "2940634",
      "postDate": "07/30/2024 09:09:40",
      "content": "<p>Thanks! Sure, so we used Polars for efficient reading of all the Parquet files combined with Python's ThreadPoolExecutor for multi-processing. It took around 10-15mins to read all 125k patents (locally at least). After reading all the patents into the memory we just stored them in a Dictionary keyed from the patent-application. Hopefully that makes sense, happy to share the code if you would like!</p>",
      "rawMarkdown": "Thanks! Sure, so we used Polars for efficient reading of all the Parquet files combined with Python's ThreadPoolExecutor for multi-processing. It took around 10-15mins to read all 125k patents (locally at least). After reading all the patents into the memory we just stored them in a Dictionary keyed from the patent-application. Hopefully that makes sense, happy to share the code if you would like!",
      "votes": null
    },
    {
      "id": "2940652",
      "postDate": "07/30/2024 09:37:55",
      "content": "<p>I see. That's why I failed all the way as I tried on Kaggle. Can you share your HW spec? </p>",
      "rawMarkdown": "I see. That's why I failed all the way as I tried on Kaggle. Can you share your HW spec?",
      "votes": null
    },
    {
      "id": "2941517",
      "postDate": "07/31/2024 04:07:37",
      "content": "<p>8 core CPU with 32GB ram, and i was using the GPU T4x2 notebook in Kaggle</p>",
      "rawMarkdown": "8 core CPU with 32GB ram, and i was using the GPU T4x2 notebook in Kaggle",
      "votes": null
    },
    {
      "id": "2941630",
      "postDate": "07/31/2024 06:44:49",
      "content": "<p>I'll try out your way - thank you so much! </p>",
      "rawMarkdown": "I'll try out your way - thank you so much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2936369,
      "author_name": "aaditshukla",
      "author_url": "",
      "post_date": "07/26/2024 05:05:55",
      "content": "<p>Thank you for sharing your detailed solution and innovative approach—it's insightful how you utilized rare CPC codes and keywords to refine your search space. Congratulations on your impressive performance and contributions to the competition! <a href=\"https://www.kaggle.com/gagankaur16\" target=\"_blank\">@gagankaur16</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2936398,
      "author_name": "gowillgo",
      "author_url": "",
      "post_date": "07/26/2024 06:09:48",
      "content": "<p><a href=\"https://www.kaggle.com/gagankaur16\" target=\"_blank\">@gagankaur16</a> Thanks for sharing your approach, and congrats on your medal! Can you explain how you did instant retrieval with Polars?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2940634,
          "author_name": "gagankaur16",
          "author_url": "",
          "post_date": "07/30/2024 09:09:40",
          "content": "<p>Thanks! Sure, so we used Polars for efficient reading of all the Parquet files combined with Python's ThreadPoolExecutor for multi-processing. It took around 10-15mins to read all 125k patents (locally at least). After reading all the patents into the memory we just stored them in a Dictionary keyed from the patent-application. Hopefully that makes sense, happy to share the code if you would like!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2940652,
              "author_name": "gowillgo",
              "author_url": "",
              "post_date": "07/30/2024 09:37:55",
              "content": "<p>I see. That's why I failed all the way as I tried on Kaggle. Can you share your HW spec? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2941517,
                  "author_name": "gagankaur16",
                  "author_url": "",
                  "post_date": "07/31/2024 04:07:37",
                  "content": "<p>8 core CPU with 32GB ram, and i was using the GPU T4x2 notebook in Kaggle</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2941630,
                      "author_name": "gowillgo",
                      "author_url": "",
                      "post_date": "07/31/2024 06:44:49",
                      "content": "<p>I'll try out your way - thank you so much! </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2936366": "Thank you to USPTO for organising and hosting this competition, and a big congratulations to all the winning teams.\n\n## Solution Overview\n\nThe overall approach was to construct a query in the form: “rare_cpcs_docs OR ((sub-query) AND (sub-query) AND …)”. We made use of a global frequency index for CPCs and an IDF lookup table for Abstract & Claims, which is discussed in detail below.\n\n- We discovered early on that just using CPC codes was very powerful for retrieval. To reduce our search space, we constructed a ‘rare_cpcs_docs’ query, which is an OR-wise concatenation of CPCs. This was only for documents that had very few global hits on CPCs (< 50). Testing locally, this was always around 5-10 documents.\n- The aim for each ‘subquery’ is to obtain common concepts to cover the remaining ~40 documents. We did this through a combination of using CPC and keywords extracted from the Abstract & Claims using KeyBERT.\n\nWe focused on rare CPC & rare keywords that shared mostly within 50 docs (i.e CPC with high local hit and low global count, keywords with high local hit and high global IDF), then we greedily picked the best performing entity until all documents were covered (Greedy Set Cover).\n\n- This process was repeated to construct as many iterations of a ‘sub-query’ as possible until we ran over the 50 token limit (no ‘magic’ token limit)\n\n**Generating ‘sub-query’**\n\nWe performed a lot of experimentation in generating these sub-queries, but none led to huge gains:\n\n- Using the Title field\n- Filtering out overlapping words across sub-queries, e.g., if “electrode” is used in sub-query-1, filter out “negative electrode” so it’s not used in subsequent queries\n- ADJ/NEAR within ngram (>=3gram) gives more choices of rare keywords that have higher hit within 50 documents. However, it costs more query term than other operators.\n\n**Constructing Global Indexes**\n\nWe analysed the entire set of parquet files (10+ million individual patents) to construct these indexes:\n\n- Frequency count of CPC codes\n- Unigram & bigram IDF values for Abstract\n- Unigram & bigram IDF values for Claims\n\nThe aim was to use these lookup tables to penalise high frequency terms, thus reducing the number of extraneous documents we retrieved.\n\n### Other Ideas\n\n- Our first entry consisted entirely of CPC codes in the form of “cpc:(a OR b OR c OR …)”. The codes were chosen based on the factor of local frequency, i.e., we prioritised codes that covered most of the 50 documents first; a simple greedy set cover algorithm. This simple query scored 0.37 on the leaderboard.\n- We also tried to focus the ‘subquery’ on each part of a patent such as (ti:subquery) (clm:subquery)(cpc:subquery). This would help to detect the common patent landscape of each part over 50 docs. Due to the small query size on each subquery, this solution didn’t produce the expected improvements.\n\n### Tips & Tricks\n\n- We used Polars + multi-threading to very quickly process and store the 125,000 relevant patents in memory for instant retrieval",
    "2936369": "Thank you for sharing your detailed solution and innovative approach—it's insightful how you utilized rare CPC codes and keywords to refine your search space. Congratulations on your impressive performance and contributions to the competition! @gagankaur16",
    "2936398": "gagankaur16 Thanks for sharing your approach, and congrats on your medal! Can you explain how you did instant retrieval with Polars?",
    "2940634": "Thanks! Sure, so we used Polars for efficient reading of all the Parquet files combined with Python's ThreadPoolExecutor for multi-processing. It took around 10-15mins to read all 125k patents (locally at least). After reading all the patents into the memory we just stored them in a Dictionary keyed from the patent-application. Hopefully that makes sense, happy to share the code if you would like!",
    "2940652": "I see. That's why I failed all the way as I tried on Kaggle. Can you share your HW spec?",
    "2941517": "8 core CPU with 32GB ram, and i was using the GPU T4x2 notebook in Kaggle",
    "2941630": "I'll try out your way - thank you so much!"
  },
  "source": "meta"
}