{
  "id": 498599,
  "title": "What the competition is asking (thinking out loud)",
  "url": "/competitions/uspto-explainable-ai/discussion/498599",
  "author_name": "",
  "post_date": "2024-04-28T23:57:09.659356300Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I'm trying to wrap my head around this competition. I think I understand what is being asked, but am not sure. The picture I have in my head is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F557179%2F9ca5905303ebb89800b89188de7deefa%2Fuspto.png?generation=1714348438706490&amp;alt=media\"></p>\n<p>Is this a possible way to approach the problem? The part that confuses me the most is what to make of the whoosh index. In the data explanation it states, </p>\n<blockquote>\n  <p>The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook. </p>\n</blockquote>\n<p>Does this imply that there will be an index given that is outside of our control, or are we supposed to supply the index that we created when running the submission?</p>\n<p>Thanks for help in better understanding what is being asked. </p>",
  "messages": [
    {
      "id": "2781679",
      "postDate": "04/28/2024 23:57:09",
      "content": "<p>Hi all,</p>\n<p>I'm trying to wrap my head around this competition. I think I understand what is being asked, but am not sure. The picture I have in my head is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F557179%2F9ca5905303ebb89800b89188de7deefa%2Fuspto.png?generation=1714348438706490&amp;alt=media\"></p>\n<p>Is this a possible way to approach the problem? The part that confuses me the most is what to make of the whoosh index. In the data explanation it states, </p>\n<blockquote>\n  <p>The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook. </p>\n</blockquote>\n<p>Does this imply that there will be an index given that is outside of our control, or are we supposed to supply the index that we created when running the submission?</p>\n<p>Thanks for help in better understanding what is being asked. </p>",
      "rawMarkdown": "Hi all,\n\nI'm trying to wrap my head around this competition. I think I understand what is being asked, but am not sure. The picture I have in my head is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F557179%2F9ca5905303ebb89800b89188de7deefa%2Fuspto.png?generation=1714348438706490&alt=media)\n\nIs this a possible way to approach the problem? The part that confuses me the most is what to make of the whoosh index. In the data explanation it states, \n\n>The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook. \n\nDoes this imply that there will be an index given that is outside of our control, or are we supposed to supply the index that we created when running the submission?\n\nThanks for help in better understanding what is being asked.",
      "votes": null
    },
    {
      "id": "2782178",
      "postDate": "04/29/2024 07:19:40",
      "content": "<p>From what I understand, the Whoosh Index helps you validate your queries during / after training. I assume that all patents, even the submission patents, are in the index - so you maybe can use features from the index for your queries too.</p>\n<p>I think you are missing the patent data that is in the parquet files. These files are containing probably the most relevant information, in terms of ways too match them etc. As far as I know you cannot access the parquet file information from a whosh index; but please correct me as this probably reduces the additional memory footprint a lot.</p>",
      "rawMarkdown": "From what I understand, the Whoosh Index helps you validate your queries during / after training. I assume that all patents, even the submission patents, are in the index - so you maybe can use features from the index for your queries too.\n\nI think you are missing the patent data that is in the parquet files. These files are containing probably the most relevant information, in terms of ways too match them etc. As far as I know you cannot access the parquet file information from a whosh index; but please correct me as this probably reduces the additional memory footprint a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2782178,
      "author_name": "valentinwerner",
      "author_url": "",
      "post_date": "04/29/2024 07:19:40",
      "content": "<p>From what I understand, the Whoosh Index helps you validate your queries during / after training. I assume that all patents, even the submission patents, are in the index - so you maybe can use features from the index for your queries too.</p>\n<p>I think you are missing the patent data that is in the parquet files. These files are containing probably the most relevant information, in terms of ways too match them etc. As far as I know you cannot access the parquet file information from a whosh index; but please correct me as this probably reduces the additional memory footprint a lot.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2781679": "Hi all,\n\nI'm trying to wrap my head around this competition. I think I understand what is being asked, but am not sure. The picture I have in my head is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F557179%2F9ca5905303ebb89800b89188de7deefa%2Fuspto.png?generation=1714348438706490&alt=media)\n\nIs this a possible way to approach the problem? The part that confuses me the most is what to make of the whoosh index. In the data explanation it states, \n\n>The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook. \n\nDoes this imply that there will be an index given that is outside of our control, or are we supposed to supply the index that we created when running the submission?\n\nThanks for help in better understanding what is being asked.",
    "2782178": "From what I understand, the Whoosh Index helps you validate your queries during / after training. I assume that all patents, even the submission patents, are in the index - so you maybe can use features from the index for your queries too.\n\nI think you are missing the patent data that is in the parquet files. These files are containing probably the most relevant information, in terms of ways too match them etc. As far as I know you cannot access the parquet file information from a whosh index; but please correct me as this probably reduces the additional memory footprint a lot."
  },
  "source": "meta"
}