{
  "id": 522208,
  "title": "10th Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/cy-10th-place-solution",
  "author_name": "",
  "post_date": "2024-07-25T01:19:54.143Z",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, congratulations to everyone who won, got a prize, medals, etc…!<br>\nThough we could not become aware of magics like \"AND\" related tokens as 1 token, we would like to share a brief solution.</p>\n<h3>1. Creating (cpc, ti, ab, detd) Pair Related to Target Patents, and Not Other Patents in Test</h3>\n<p>First, we calculated all cpc_codes, titles, etc., related to each test rows, and generated pairs. If each token pair was not related to other test patents and not more related to other patents not included in test.csv than patents in test rows, we used this pair.</p>\n<h3>2. Sorting Pairs To Save Token Counts</h3>\n<p>Imagine a query like below(token counts: 14)</p>\n<pre><code>( cpc) OR ( cpc) OR ( cpc) OR ( cpc) OR ( cpc)\n</code></pre>\n<p>By merging pairs which have same token, token counts were saved (14 -&gt; 11)</p>\n<pre><code>( ( OR cpc OR cpc)) OR ( ( OR cpc))\n</code></pre>\n<h3>3. Final Thought</h3>\n<p>I was struggling with a reliable validation set like below discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/518461\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/518461</a><br>\nBy giving up creating another validation (or test) whoosh index and calculating all patent metadata, everything went fine.</p>\n<h3>4. Code Link</h3>\n<p>Notebook Link(LB: 0.83): <a href=\"https://www.kaggle.com/code/shigeria/uspto-final-sub?scriptVersionId=189656420\" target=\"_blank\">https://www.kaggle.com/code/shigeria/uspto-final-sub?scriptVersionId=189656420</a></p>",
  "messages": [
    {
      "id": "2935092",
      "postDate": "07/25/2024 01:07:51",
      "content": "<p>First of all, congratulations to everyone who won, got a prize, medals, etc…!<br>\nThough we could not become aware of magics like \"AND\" related tokens as 1 token, we would like to share a brief solution.</p>\n<h3>1. Creating (cpc, ti, ab, detd) Pair Related to Target Patents, and Not Other Patents in Test</h3>\n<p>First, we calculated all cpc_codes, titles, etc., related to each test rows, and generated pairs. If each token pair was not related to other test patents and not more related to other patents not included in test.csv than patents in test rows, we used this pair.</p>\n<h3>2. Sorting Pairs To Save Token Counts</h3>\n<p>Imagine a query like below(token counts: 14)</p>\n<pre><code>( cpc) OR ( cpc) OR ( cpc) OR ( cpc) OR ( cpc)\n</code></pre>\n<p>By merging pairs which have same token, token counts were saved (14 -&gt; 11)</p>\n<pre><code>( ( OR cpc OR cpc)) OR ( ( OR cpc))\n</code></pre>\n<h3>3. Final Thought</h3>\n<p>I was struggling with a reliable validation set like below discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/518461\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/518461</a><br>\nBy giving up creating another validation (or test) whoosh index and calculating all patent metadata, everything went fine.</p>\n<h3>4. Code Link</h3>\n<p>Notebook Link(LB: 0.83): <a href=\"https://www.kaggle.com/code/shigeria/uspto-final-sub?scriptVersionId=189656420\" target=\"_blank\">https://www.kaggle.com/code/shigeria/uspto-final-sub?scriptVersionId=189656420</a></p>",
      "rawMarkdown": "First of all, congratulations to everyone who won, got a prize, medals, etc...!\nThough we could not become aware of magics like \"AND\" related tokens as 1 token, we would like to share a brief solution.\n\n### 1. Creating (cpc, ti, ab, detd) Pair Related to Target Patents, and Not Other Patents in Test\nFirst, we calculated all cpc_codes, titles, etc., related to each test rows, and generated pairs. If each token pair was not related to other test patents and not more related to other patents not included in test.csv than patents in test rows, we used this pair.\n\n### 2. Sorting Pairs To Save Token Counts\nImagine a query like below(token counts: 14)\n```\n(cpc:aaa cpc:bbb) OR (cpc:aaa cpc:ccc) OR (cpc:aaa cpc:ddd) OR (cpc:bbb cpc:ddd) OR (cpc:bbb cpc:eee)\n```\nBy merging pairs which have same token, token counts were saved (14 -> 11)\n```\n(cpc:aaa (cpc:bbb OR cpc:ccc OR cpc:ddd)) OR (cpc:bbb (cpc:ddd OR cpc:eee))\n```\n### 3. Final Thought\nI was struggling with a reliable validation set like below discussion.\nhttps://www.kaggle.com/competitions/uspto-explainable-ai/discussion/518461\nBy giving up creating another validation (or test) whoosh index and calculating all patent metadata, everything went fine.\n\n###4. Code Link\nNotebook Link(LB: 0.83): https://www.kaggle.com/code/shigeria/uspto-final-sub?scriptVersionId=189656420",
      "votes": null
    },
    {
      "id": "2935301",
      "postDate": "07/25/2024 06:34:09",
      "content": "<p>Congratulations on winning the 10th place in this competition. Very interesting ideas indeed. </p>",
      "rawMarkdown": "Congratulations on winning the 10th place in this competition. Very interesting ideas indeed.",
      "votes": null
    },
    {
      "id": "2936716",
      "postDate": "07/26/2024 12:02:51",
      "content": "<p>Congratulations🎉</p>",
      "rawMarkdown": "Congratulations🎉",
      "votes": null
    },
    {
      "id": "2937249",
      "postDate": "07/26/2024 21:49:26",
      "content": "<p>Thanks, Snorf. Also, congratulations for your 1st gold medal in PII Data Detection!</p>",
      "rawMarkdown": "Thanks, Snorf. Also, congratulations for your 1st gold medal in PII Data Detection!",
      "votes": null
    },
    {
      "id": "2940845",
      "postDate": "07/30/2024 13:55:24",
      "content": "<p>Congratulations on your impressive achievement in the USPTO Explainable AI for Patent Professionals competition, <a href=\"https://www.kaggle.com/shigeria\" target=\"_blank\">@shigeria</a>!<br>\nSecuring 10th place with your solution is a notable accomplishment. Your approach to creating and sorting pairs of tokens related to target patents demonstrates a strategic method for optimizing query efficiency.</p>\n<p>Your solution’s focus on saving token counts by merging pairs and improving query structure highlights your attention to detail and innovative problem-solving skills. I appreciate your transparency in discussing the challenges you faced with validation and the steps you took to overcome them. Your shared code and insights will undoubtedly be valuable to others working on similar tasks. Thanks for contributing to the community with your thorough and effective approach.</p>",
      "rawMarkdown": "Congratulations on your impressive achievement in the USPTO Explainable AI for Patent Professionals competition, @shigeria!\nSecuring 10th place with your solution is a notable accomplishment. Your approach to creating and sorting pairs of tokens related to target patents demonstrates a strategic method for optimizing query efficiency.\n\nYour solution’s focus on saving token counts by merging pairs and improving query structure highlights your attention to detail and innovative problem-solving skills. I appreciate your transparency in discussing the challenges you faced with validation and the steps you took to overcome them. Your shared code and insights will undoubtedly be valuable to others working on similar tasks. Thanks for contributing to the community with your thorough and effective approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2935301,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/25/2024 06:34:09",
      "content": "<p>Congratulations on winning the 10th place in this competition. Very interesting ideas indeed. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2936716,
      "author_name": "snorfyang",
      "author_url": "",
      "post_date": "07/26/2024 12:02:51",
      "content": "<p>Congratulations🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 2937249,
          "author_name": "shigeria",
          "author_url": "",
          "post_date": "07/26/2024 21:49:26",
          "content": "<p>Thanks, Snorf. Also, congratulations for your 1st gold medal in PII Data Detection!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2940845,
      "author_name": "",
      "author_url": "",
      "post_date": "07/30/2024 13:55:24",
      "content": "<p>Congratulations on your impressive achievement in the USPTO Explainable AI for Patent Professionals competition, <a href=\"https://www.kaggle.com/shigeria\" target=\"_blank\">@shigeria</a>!<br>\nSecuring 10th place with your solution is a notable accomplishment. Your approach to creating and sorting pairs of tokens related to target patents demonstrates a strategic method for optimizing query efficiency.</p>\n<p>Your solution’s focus on saving token counts by merging pairs and improving query structure highlights your attention to detail and innovative problem-solving skills. I appreciate your transparency in discussing the challenges you faced with validation and the steps you took to overcome them. Your shared code and insights will undoubtedly be valuable to others working on similar tasks. Thanks for contributing to the community with your thorough and effective approach.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935092": "First of all, congratulations to everyone who won, got a prize, medals, etc...!\nThough we could not become aware of magics like \"AND\" related tokens as 1 token, we would like to share a brief solution.\n\n### 1. Creating (cpc, ti, ab, detd) Pair Related to Target Patents, and Not Other Patents in Test\nFirst, we calculated all cpc_codes, titles, etc., related to each test rows, and generated pairs. If each token pair was not related to other test patents and not more related to other patents not included in test.csv than patents in test rows, we used this pair.\n\n### 2. Sorting Pairs To Save Token Counts\nImagine a query like below(token counts: 14)\n```\n(cpc:aaa cpc:bbb) OR (cpc:aaa cpc:ccc) OR (cpc:aaa cpc:ddd) OR (cpc:bbb cpc:ddd) OR (cpc:bbb cpc:eee)\n```\nBy merging pairs which have same token, token counts were saved (14 -> 11)\n```\n(cpc:aaa (cpc:bbb OR cpc:ccc OR cpc:ddd)) OR (cpc:bbb (cpc:ddd OR cpc:eee))\n```\n### 3. Final Thought\nI was struggling with a reliable validation set like below discussion.\nhttps://www.kaggle.com/competitions/uspto-explainable-ai/discussion/518461\nBy giving up creating another validation (or test) whoosh index and calculating all patent metadata, everything went fine.\n\n###4. Code Link\nNotebook Link(LB: 0.83): https://www.kaggle.com/code/shigeria/uspto-final-sub?scriptVersionId=189656420",
    "2935301": "Congratulations on winning the 10th place in this competition. Very interesting ideas indeed.",
    "2936716": "Congratulations🎉",
    "2937249": "Thanks, Snorf. Also, congratulations for your 1st gold medal in PII Data Detection!",
    "2940845": "Congratulations on your impressive achievement in the USPTO Explainable AI for Patent Professionals competition, @shigeria!\nSecuring 10th place with your solution is a notable accomplishment. Your approach to creating and sorting pairs of tokens related to target patents demonstrates a strategic method for optimizing query efficiency.\n\nYour solution’s focus on saving token counts by merging pairs and improving query structure highlights your attention to detail and innovative problem-solving skills. I appreciate your transparency in discussing the challenges you faced with validation and the steps you took to overcome them. Your shared code and insights will undoubtedly be valuable to others working on similar tasks. Thanks for contributing to the community with your thorough and effective approach."
  },
  "source": "meta"
}