{
  "id": 518556,
  "title": "Publication number in nearest_neighbors.csv but not in patent_data/parquets",
  "url": "/competitions/uspto-explainable-ai/discussion/518556",
  "author_name": "",
  "post_date": "2024-07-07T04:20:41.218077100Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Problem</h1>\n<p>I ran into a problem when I was collecting 2500 patents for a personal test.csv to check the runtime of my notebook, by randomly selecting publication numbers from nearest_neighbors.csv. But I got an error when one publication number was in my test.csv, was found in patent_metadata, but by publication date was not in the right patent_data (by publication date year_month). And in the end, this number was not found in any patent_data parquet (I've checked every parquet)</p>\n<h1>Demo notebook</h1>\n<p>I will attach a notebook for demonstration.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/qurusx/hidden-patent</a></p>\n<h1>Question</h1>\n<p>Maybe I'm missing or don't understand something, but is it supposed to be like this?</p>",
  "messages": [
    {
      "id": "2909508",
      "postDate": "07/07/2024 04:20:41",
      "content": "<h1>Problem</h1>\n<p>I ran into a problem when I was collecting 2500 patents for a personal test.csv to check the runtime of my notebook, by randomly selecting publication numbers from nearest_neighbors.csv. But I got an error when one publication number was in my test.csv, was found in patent_metadata, but by publication date was not in the right patent_data (by publication date year_month). And in the end, this number was not found in any patent_data parquet (I've checked every parquet)</p>\n<h1>Demo notebook</h1>\n<p>I will attach a notebook for demonstration.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/qurusx/hidden-patent</a></p>\n<h1>Question</h1>\n<p>Maybe I'm missing or don't understand something, but is it supposed to be like this?</p>",
      "rawMarkdown": "# Problem\nI ran into a problem when I was collecting 2500 patents for a personal test.csv to check the runtime of my notebook, by randomly selecting publication numbers from nearest_neighbors.csv. But I got an error when one publication number was in my test.csv, was found in patent_metadata, but by publication date was not in the right patent_data (by publication date year_month). And in the end, this number was not found in any patent_data parquet (I've checked every parquet)\n# Demo notebook\nI will attach a notebook for demonstration.\n[https://www.kaggle.com/code/qurusx/hidden-patent](url)\n# Question\nMaybe I'm missing or don't understand something, but is it supposed to be like this?",
      "votes": null
    },
    {
      "id": "2916668",
      "postDate": "07/11/2024 07:06:05",
      "content": "<p>Yes there are some publication numbers that have metadata but not raw field data, here is a notebook listing them: <a href=\"https://www.kaggle.com/dilliontan/uspto-metadata-raw-data-check\" target=\"_blank\">https://www.kaggle.com/dilliontan/uspto-metadata-raw-data-check</a></p>\n<p>I assume this is part of the competition challenge, a very small number of patents don't have raw data, I don't think this affects the final scores much. But it can cause notebook submission errors (Notebook Threw Exception) if just fetching based on publication date without checking existence</p>",
      "rawMarkdown": "Yes there are some publication numbers that have metadata but not raw field data, here is a notebook listing them: https://www.kaggle.com/dilliontan/uspto-metadata-raw-data-check\n\nI assume this is part of the competition challenge, a very small number of patents don't have raw data, I don't think this affects the final scores much. But it can cause notebook submission errors (Notebook Threw Exception) if just fetching based on publication date without checking existence",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2916668,
      "author_name": "dilliontan",
      "author_url": "",
      "post_date": "07/11/2024 07:06:05",
      "content": "<p>Yes there are some publication numbers that have metadata but not raw field data, here is a notebook listing them: <a href=\"https://www.kaggle.com/dilliontan/uspto-metadata-raw-data-check\" target=\"_blank\">https://www.kaggle.com/dilliontan/uspto-metadata-raw-data-check</a></p>\n<p>I assume this is part of the competition challenge, a very small number of patents don't have raw data, I don't think this affects the final scores much. But it can cause notebook submission errors (Notebook Threw Exception) if just fetching based on publication date without checking existence</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2909508": "# Problem\nI ran into a problem when I was collecting 2500 patents for a personal test.csv to check the runtime of my notebook, by randomly selecting publication numbers from nearest_neighbors.csv. But I got an error when one publication number was in my test.csv, was found in patent_metadata, but by publication date was not in the right patent_data (by publication date year_month). And in the end, this number was not found in any patent_data parquet (I've checked every parquet)\n# Demo notebook\nI will attach a notebook for demonstration.\n[https://www.kaggle.com/code/qurusx/hidden-patent](url)\n# Question\nMaybe I'm missing or don't understand something, but is it supposed to be like this?",
    "2916668": "Yes there are some publication numbers that have metadata but not raw field data, here is a notebook listing them: https://www.kaggle.com/dilliontan/uspto-metadata-raw-data-check\n\nI assume this is part of the competition challenge, a very small number of patents don't have raw data, I don't think this affects the final scores much. But it can cause notebook submission errors (Notebook Threw Exception) if just fetching based on publication date without checking existence"
  },
  "source": "meta"
}