{
  "id": 497556,
  "title": "Nearest Neighbors based on patent embeddings (vector size 64)",
  "url": "/competitions/uspto-explainable-ai/discussion/497556",
  "author_name": "",
  "post_date": "2024-04-25T02:46:14.136273300Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F082491286038d2427f3e3f338f876e5b%2FScreenshot%202024-04-27%20at%205.04.23AM.png?generation=1714174499947051&amp;alt=media\"></p>\n<blockquote>\n  <p><strong>64 scalar vector =&gt; (Version 1) Machine-learned vector embedding based on document contents and metadata, where two documents that have similar technical content have a high dot product score of their embedding vectors.</strong></p>\n</blockquote>\n<hr>\n<p><a href=\"https://cloud.google.com/blog/products/data-analytics/expanding-your-patent-set-with-ml-and-bigquery\" target=\"_blank\">Source</a></p>\n<blockquote>\n  <p>The <strong>patent embeddings were built using a machine learning model that predicted a patent's CPC code from its text</strong>. Therefore, the <strong>learned embeddings are a vector of 64</strong> continuous numbers intended to encode the information in a patent's text. Distances between the embeddings can then be calculated and used as a measure of similarity between two patents.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>So, we are reverse engineering the vector to search query in this competition?</strong></p>\n</blockquote>\n<hr>\n<pre><code>             \n</code></pre>\n<p><a href=\"https://www.kaggle.com/code/jayyonamine/using-bigquery-for-patent-analysis\" target=\"_blank\"><strong>Notebook - BigQuery for Patents</strong></a></p>",
  "messages": [
    {
      "id": "2773996",
      "postDate": "04/25/2024 02:46:14",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F082491286038d2427f3e3f338f876e5b%2FScreenshot%202024-04-27%20at%205.04.23AM.png?generation=1714174499947051&amp;alt=media\"></p>\n<blockquote>\n  <p><strong>64 scalar vector =&gt; (Version 1) Machine-learned vector embedding based on document contents and metadata, where two documents that have similar technical content have a high dot product score of their embedding vectors.</strong></p>\n</blockquote>\n<hr>\n<p><a href=\"https://cloud.google.com/blog/products/data-analytics/expanding-your-patent-set-with-ml-and-bigquery\" target=\"_blank\">Source</a></p>\n<blockquote>\n  <p>The <strong>patent embeddings were built using a machine learning model that predicted a patent's CPC code from its text</strong>. Therefore, the <strong>learned embeddings are a vector of 64</strong> continuous numbers intended to encode the information in a patent's text. Distances between the embeddings can then be calculated and used as a measure of similarity between two patents.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>So, we are reverse engineering the vector to search query in this competition?</strong></p>\n</blockquote>\n<hr>\n<pre><code>             \n</code></pre>\n<p><a href=\"https://www.kaggle.com/code/jayyonamine/using-bigquery-for-patent-analysis\" target=\"_blank\"><strong>Notebook - BigQuery for Patents</strong></a></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F082491286038d2427f3e3f338f876e5b%2FScreenshot%202024-04-27%20at%205.04.23AM.png?generation=1714174499947051&alt=media)\n> **64 scalar vector => (Version 1) Machine-learned vector embedding based on document contents and metadata, where two documents that have similar technical content have a high dot product score of their embedding vectors.**\n\n---\n\n[Source](https://cloud.google.com/blog/products/data-analytics/expanding-your-patent-set-with-ml-and-bigquery)\n> The **patent embeddings were built using a machine learning model that predicted a patent's CPC code from its text**. Therefore, the **learned embeddings are a vector of 64** continuous numbers intended to encode the information in a patent's text. Distances between the embeddings can then be calculated and used as a measure of similarity between two patents.\n\n---\n\n> **So, we are reverse engineering the vector to search query in this competition?**\n\n---\n\n```yaml\ntest.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n```\n\n\n[**Notebook - BigQuery for Patents**](https://www.kaggle.com/code/jayyonamine/using-bigquery-for-patent-analysis)",
      "votes": null
    },
    {
      "id": "2774859",
      "postDate": "04/25/2024 10:44:39",
      "content": "<p>Yes, the goal is to explain the similarity results as structured search terms, not to build a better similarity model.</p>",
      "rawMarkdown": "Yes, the goal is to explain the similarity results as structured search terms, not to build a better similarity model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2774859,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "04/25/2024 10:44:39",
      "content": "<p>Yes, the goal is to explain the similarity results as structured search terms, not to build a better similarity model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2773996": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F082491286038d2427f3e3f338f876e5b%2FScreenshot%202024-04-27%20at%205.04.23AM.png?generation=1714174499947051&alt=media)\n> **64 scalar vector => (Version 1) Machine-learned vector embedding based on document contents and metadata, where two documents that have similar technical content have a high dot product score of their embedding vectors.**\n\n---\n\n[Source](https://cloud.google.com/blog/products/data-analytics/expanding-your-patent-set-with-ml-and-bigquery)\n> The **patent embeddings were built using a machine learning model that predicted a patent's CPC code from its text**. Therefore, the **learned embeddings are a vector of 64** continuous numbers intended to encode the information in a patent's text. Distances between the embeddings can then be calculated and used as a measure of similarity between two patents.\n\n---\n\n> **So, we are reverse engineering the vector to search query in this competition?**\n\n---\n\n```yaml\ntest.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n```\n\n\n[**Notebook - BigQuery for Patents**](https://www.kaggle.com/code/jayyonamine/using-bigquery-for-patent-analysis)",
    "2774859": "Yes, the goal is to explain the similarity results as structured search terms, not to build a better similarity model."
  },
  "source": "meta"
}