{
  "id": 497593,
  "title": "Initial thoughts - ngrams and overlap",
  "url": "/competitions/uspto-explainable-ai/discussion/497593",
  "author_name": "",
  "post_date": "2024-04-25T05:39:19.602444200Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This is an interesting competition!</p>\n<p>If I understand this competition correctly, for each sample, we will be given 50 patents that have been grouped together based on Google's embeddings, and then we need to create a query that can pull up these same results.</p>\n<p>My initial idea is to find common unigrams or bigrams in each of the patent's section (title, abstract, claims, description) and use that in an AND query. This requires no training and is relatively straightforward to do.</p>\n<p>I next thought about how to utilize the NOT of the query. For this, I think an embedding approach could be used to find similar patents that are just outside the range of the top 50. Common ngrams in these documents that are not in the closest 50 would then be NOT terms in the query.</p>\n<p>Since this is focused on interpretability, I'm wondering how useful LLMs would be since the outputs can be somewhat unpredictable and uninterpretable.</p>\n<p>What are your thoughts?</p>",
  "messages": [
    {
      "id": "2774262",
      "postDate": "04/25/2024 05:39:19",
      "content": "<p>This is an interesting competition!</p>\n<p>If I understand this competition correctly, for each sample, we will be given 50 patents that have been grouped together based on Google's embeddings, and then we need to create a query that can pull up these same results.</p>\n<p>My initial idea is to find common unigrams or bigrams in each of the patent's section (title, abstract, claims, description) and use that in an AND query. This requires no training and is relatively straightforward to do.</p>\n<p>I next thought about how to utilize the NOT of the query. For this, I think an embedding approach could be used to find similar patents that are just outside the range of the top 50. Common ngrams in these documents that are not in the closest 50 would then be NOT terms in the query.</p>\n<p>Since this is focused on interpretability, I'm wondering how useful LLMs would be since the outputs can be somewhat unpredictable and uninterpretable.</p>\n<p>What are your thoughts?</p>",
      "rawMarkdown": "This is an interesting competition!\n\nIf I understand this competition correctly, for each sample, we will be given 50 patents that have been grouped together based on Google's embeddings, and then we need to create a query that can pull up these same results.\n\nMy initial idea is to find common unigrams or bigrams in each of the patent's section (title, abstract, claims, description) and use that in an AND query. This requires no training and is relatively straightforward to do.\n\nI next thought about how to utilize the NOT of the query. For this, I think an embedding approach could be used to find similar patents that are just outside the range of the top 50. Common ngrams in these documents that are not in the closest 50 would then be NOT terms in the query.\n\n\nSince this is focused on interpretability, I'm wondering how useful LLMs would be since the outputs can be somewhat unpredictable and uninterpretable.\n\nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "2776186",
      "postDate": "04/26/2024 03:42:58",
      "content": "<p>I think this is one of the most fascinating competitions I've seen in Kaggle. I'm ready to join.🤘</p>",
      "rawMarkdown": "I think this is one of the most fascinating competitions I've seen in Kaggle. I'm ready to join.🤘",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2776186,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "04/26/2024 03:42:58",
      "content": "<p>I think this is one of the most fascinating competitions I've seen in Kaggle. I'm ready to join.🤘</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2774262": "This is an interesting competition!\n\nIf I understand this competition correctly, for each sample, we will be given 50 patents that have been grouped together based on Google's embeddings, and then we need to create a query that can pull up these same results.\n\nMy initial idea is to find common unigrams or bigrams in each of the patent's section (title, abstract, claims, description) and use that in an AND query. This requires no training and is relatively straightforward to do.\n\nI next thought about how to utilize the NOT of the query. For this, I think an embedding approach could be used to find similar patents that are just outside the range of the top 50. Common ngrams in these documents that are not in the closest 50 would then be NOT terms in the query.\n\n\nSince this is focused on interpretability, I'm wondering how useful LLMs would be since the outputs can be somewhat unpredictable and uninterpretable.\n\nWhat are your thoughts?",
    "2776186": "I think this is one of the most fascinating competitions I've seen in Kaggle. I'm ready to join.🤘"
  },
  "source": "meta"
}