{
  "id": 522207,
  "title": "11th Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/devin-11th-place-solution",
  "author_name": "",
  "post_date": "2024-07-25T01:07:19.443Z",
  "votes": 21,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting the competition and thanks to everyone who contributed code and ideas. </p>\n<h2>Overview</h2>\n<ul>\n<li>Queries consisted of infrequent tokens connected by <code>OR</code></li>\n<li>Used single words and <code>ADJ</code> pairs</li>\n<li>No magic, and simple rules to find decent token combinations for query</li>\n</ul>\n<h2>Validation</h2>\n<p>Used 2000 patents and the 100 nearest neighbors. I pulled the embedding data using BigQuery and then created dataset with 100 nearest neighbors. This just about perfectly matched the leaderboard so I wonder if the test set also used more then the 50 nearest neighbors.</p>\n<h2>Approach</h2>\n<p>Early on I realized that we can identify individual patents by very infrequent words (plenty of misspellings). Just using infrequent tokens and and <code>OR</code> gets 0.77. I created datasets containing the frequencies of all the words for all patents, and then datasets containing only the less frequent tokens. </p>\n<p>In addition to uniquely identifying patents with infrequent tokens I also identified groups of patents that shared the same infrequent token. </p>\n<p>I had to stop competing on Kaggle a couple of weeks into this competition and this is were I left my solution. </p>\n<p>In the last few days seeing that I could maybe get a gold meddle if I improved a little and I added directly <code>ADJ</code> pairs (no gap, or distance of 1) for title, abstract and claims. If I could get at least two of the target patents and the adjacent pair appeared less then 8 times in the full dataset then I would use the pair. This brought my score to 0.81. I got lucky on the shake and just squeezed into the gold section (sorry dt). </p>\n<p>My approach was not very optimized and I am excited to see some interesting solutions being posted. I didn't find any of the magic for lengthening queries, but relieved that I didn't miss anything super obvious to improve score to upper 90s.</p>",
  "messages": [
    {
      "id": "2935089",
      "postDate": "07/25/2024 01:03:21",
      "content": "<p>Thanks to the organizers for hosting the competition and thanks to everyone who contributed code and ideas. </p>\n<h2>Overview</h2>\n<ul>\n<li>Queries consisted of infrequent tokens connected by <code>OR</code></li>\n<li>Used single words and <code>ADJ</code> pairs</li>\n<li>No magic, and simple rules to find decent token combinations for query</li>\n</ul>\n<h2>Validation</h2>\n<p>Used 2000 patents and the 100 nearest neighbors. I pulled the embedding data using BigQuery and then created dataset with 100 nearest neighbors. This just about perfectly matched the leaderboard so I wonder if the test set also used more then the 50 nearest neighbors.</p>\n<h2>Approach</h2>\n<p>Early on I realized that we can identify individual patents by very infrequent words (plenty of misspellings). Just using infrequent tokens and and <code>OR</code> gets 0.77. I created datasets containing the frequencies of all the words for all patents, and then datasets containing only the less frequent tokens. </p>\n<p>In addition to uniquely identifying patents with infrequent tokens I also identified groups of patents that shared the same infrequent token. </p>\n<p>I had to stop competing on Kaggle a couple of weeks into this competition and this is were I left my solution. </p>\n<p>In the last few days seeing that I could maybe get a gold meddle if I improved a little and I added directly <code>ADJ</code> pairs (no gap, or distance of 1) for title, abstract and claims. If I could get at least two of the target patents and the adjacent pair appeared less then 8 times in the full dataset then I would use the pair. This brought my score to 0.81. I got lucky on the shake and just squeezed into the gold section (sorry dt). </p>\n<p>My approach was not very optimized and I am excited to see some interesting solutions being posted. I didn't find any of the magic for lengthening queries, but relieved that I didn't miss anything super obvious to improve score to upper 90s.</p>",
      "rawMarkdown": "Thanks to the organizers for hosting the competition and thanks to everyone who contributed code and ideas. \n\n## Overview\n- Queries consisted of infrequent tokens connected by `OR`\n- Used single words and `ADJ` pairs\n- No magic, and simple rules to find decent token combinations for query\n\n## Validation\nUsed 2000 patents and the 100 nearest neighbors. I pulled the embedding data using BigQuery and then created dataset with 100 nearest neighbors. This just about perfectly matched the leaderboard so I wonder if the test set also used more then the 50 nearest neighbors.\n\n## Approach\nEarly on I realized that we can identify individual patents by very infrequent words (plenty of misspellings). Just using infrequent tokens and and `OR` gets 0.77. I created datasets containing the frequencies of all the words for all patents, and then datasets containing only the less frequent tokens. \n\nIn addition to uniquely identifying patents with infrequent tokens I also identified groups of patents that shared the same infrequent token. \n\nI had to stop competing on Kaggle a couple of weeks into this competition and this is were I left my solution. \n\nIn the last few days seeing that I could maybe get a gold meddle if I improved a little and I added directly `ADJ` pairs (no gap, or distance of 1) for title, abstract and claims. If I could get at least two of the target patents and the adjacent pair appeared less then 8 times in the full dataset then I would use the pair. This brought my score to 0.81. I got lucky on the shake and just squeezed into the gold section (sorry dt). \n\nMy approach was not very optimized and I am excited to see some interesting solutions being posted. I didn't find any of the magic for lengthening queries, but relieved that I didn't miss anything super obvious to improve score to upper 90s.",
      "votes": null
    },
    {
      "id": "2935162",
      "postDate": "07/25/2024 03:22:06",
      "content": "<p>Congrats Devin! super sad on the shake because I really wanted a gold medal on my second competition, but you did well with so few submissions and I knew it was going to be risky at the end. Nice find on the validation set. I will post my solution soon, believe its quite different from the rest</p>",
      "rawMarkdown": "Congrats Devin! super sad on the shake because I really wanted a gold medal on my second competition, but you did well with so few submissions and I knew it was going to be risky at the end. Nice find on the validation set. I will post my solution soon, believe its quite different from the rest",
      "votes": null
    },
    {
      "id": "2935304",
      "postDate": "07/25/2024 06:37:42",
      "content": "<p>Congratulations on winning the 11th place in this competition. <br>\nThanks for sharing your experiences and insights. </p>",
      "rawMarkdown": "Congratulations on winning the 11th place in this competition. \nThanks for sharing your experiences and insights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2935162,
      "author_name": "dilliontan",
      "author_url": "",
      "post_date": "07/25/2024 03:22:06",
      "content": "<p>Congrats Devin! super sad on the shake because I really wanted a gold medal on my second competition, but you did well with so few submissions and I knew it was going to be risky at the end. Nice find on the validation set. I will post my solution soon, believe its quite different from the rest</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2935304,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/25/2024 06:37:42",
      "content": "<p>Congratulations on winning the 11th place in this competition. <br>\nThanks for sharing your experiences and insights. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935089": "Thanks to the organizers for hosting the competition and thanks to everyone who contributed code and ideas. \n\n## Overview\n- Queries consisted of infrequent tokens connected by `OR`\n- Used single words and `ADJ` pairs\n- No magic, and simple rules to find decent token combinations for query\n\n## Validation\nUsed 2000 patents and the 100 nearest neighbors. I pulled the embedding data using BigQuery and then created dataset with 100 nearest neighbors. This just about perfectly matched the leaderboard so I wonder if the test set also used more then the 50 nearest neighbors.\n\n## Approach\nEarly on I realized that we can identify individual patents by very infrequent words (plenty of misspellings). Just using infrequent tokens and and `OR` gets 0.77. I created datasets containing the frequencies of all the words for all patents, and then datasets containing only the less frequent tokens. \n\nIn addition to uniquely identifying patents with infrequent tokens I also identified groups of patents that shared the same infrequent token. \n\nI had to stop competing on Kaggle a couple of weeks into this competition and this is were I left my solution. \n\nIn the last few days seeing that I could maybe get a gold meddle if I improved a little and I added directly `ADJ` pairs (no gap, or distance of 1) for title, abstract and claims. If I could get at least two of the target patents and the adjacent pair appeared less then 8 times in the full dataset then I would use the pair. This brought my score to 0.81. I got lucky on the shake and just squeezed into the gold section (sorry dt). \n\nMy approach was not very optimized and I am excited to see some interesting solutions being posted. I didn't find any of the magic for lengthening queries, but relieved that I didn't miss anything super obvious to improve score to upper 90s.",
    "2935162": "Congrats Devin! super sad on the shake because I really wanted a gold medal on my second competition, but you did well with so few submissions and I knew it was going to be risky at the end. Nice find on the validation set. I will post my solution soon, believe its quite different from the rest",
    "2935304": "Congratulations on winning the 11th place in this competition. \nThanks for sharing your experiences and insights."
  },
  "source": "meta"
}