{
  "id": 519365,
  "title": "Find patent description in 0.1 ms - O(1) time solution",
  "url": "/competitions/uspto-explainable-ai/discussion/519365",
  "author_name": "",
  "post_date": "2024-07-10T19:45:04.930365100Z",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Problem</h2>\n<p>One of the problems with this contest is the use of competition data. If you want to extract data from a single patent, you have to access patent_data/year_month.parquet and read the data. Yes, you can use tricks like reading inline csv, using pyarrow, Rust libraries and other things. But the problem is the time and RAM of loading the information of a single patent (tens of seconds and gigabytes of memory).</p>\n<h2>Solution</h2>\n<p>The solution to the problem is to use HDF5 files to store patent information. Namely, storing patent information for n-number of years in multiple files (e.g. 2000_1.h5) and files in a notebook output (up to 20 GB). Loading these notebooks into your main code. HDF5 provide access to data as in hashmap for O(1) time, in my case by 'publication_nubmer'. Open the file - get the information of the desired patent without loading the rest of the patents.</p>\n<h2>Links to finished notebooks (11 in total)</h2>\n<p>Each notebook contains patents from year <strong>i up j</strong><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1700up1929\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1700up1929</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1930up1974\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1930up1974</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1975up1994\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1975up1994</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1995up2004\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1995up2004</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2005up2009\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2005up2009</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2010up2013\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2010up2013</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2014up2016\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2014up2016</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2017up2019\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2017up2019</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2020up2021\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2020up2021</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2022up2022\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2022up2022</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2023up2023\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2023up2023</a></p>\n<h2>A link to a notebook with an example of how to use it</h2>\n<p><a href=\"https://www.kaggle.com/code/qurusx/description-fast-extract\" target=\"_blank\">https://www.kaggle.com/code/qurusx/description-fast-extract</a></p>\n<h2>What ideas didn't work (if interested)</h2>\n<p><strong>The first one</strong><br>\nUsing pkl files. We form the parquet into a folder with the same name, and each line in it into a separate pkl file (a line can be represented as an array). <br>\nWhy no?<br>\nIf the notebook output contains more than 100,000 files - it is packed into a zip file.<br>\nOkay, we split it even smaller. Load 160 notebooks into the main code - we get an error that the notebook cannot be saved.<br>\n<strong>Second</strong><br>\nUsing your own zip files (ZIP64, tar, zip7). Folders with pkl files in zip.<br>\nWhy no?<br>\nOpening a ZIP file takes a few seconds and hundreds of megabytes of RAM.</p>\n<h2>Why it can be useful?</h2>\n<p>Most patents have a title and description rather than abstract or claims. Now you can get the description (the most complete information) of each patent for 0.1 ms from test.csv when you submit. I have my notebook running for an average of 5 hours.</p>\n<h2>Tip</h2>\n<p>There are 2 weeks left till the end of the competition, I hope this will help someone.</p>",
  "messages": [
    {
      "id": "2916129",
      "postDate": "07/10/2024 19:45:04",
      "content": "<h2>Problem</h2>\n<p>One of the problems with this contest is the use of competition data. If you want to extract data from a single patent, you have to access patent_data/year_month.parquet and read the data. Yes, you can use tricks like reading inline csv, using pyarrow, Rust libraries and other things. But the problem is the time and RAM of loading the information of a single patent (tens of seconds and gigabytes of memory).</p>\n<h2>Solution</h2>\n<p>The solution to the problem is to use HDF5 files to store patent information. Namely, storing patent information for n-number of years in multiple files (e.g. 2000_1.h5) and files in a notebook output (up to 20 GB). Loading these notebooks into your main code. HDF5 provide access to data as in hashmap for O(1) time, in my case by 'publication_nubmer'. Open the file - get the information of the desired patent without loading the rest of the patents.</p>\n<h2>Links to finished notebooks (11 in total)</h2>\n<p>Each notebook contains patents from year <strong>i up j</strong><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1700up1929\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1700up1929</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1930up1974\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1930up1974</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1975up1994\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1975up1994</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/1995up2004\" target=\"_blank\">https://www.kaggle.com/code/qurusx/1995up2004</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2005up2009\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2005up2009</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2010up2013\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2010up2013</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2014up2016\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2014up2016</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2017up2019\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2017up2019</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2020up2021\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2020up2021</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2022up2022\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2022up2022</a><br>\n<a href=\"https://www.kaggle.com/code/qurusx/2023up2023\" target=\"_blank\">https://www.kaggle.com/code/qurusx/2023up2023</a></p>\n<h2>A link to a notebook with an example of how to use it</h2>\n<p><a href=\"https://www.kaggle.com/code/qurusx/description-fast-extract\" target=\"_blank\">https://www.kaggle.com/code/qurusx/description-fast-extract</a></p>\n<h2>What ideas didn't work (if interested)</h2>\n<p><strong>The first one</strong><br>\nUsing pkl files. We form the parquet into a folder with the same name, and each line in it into a separate pkl file (a line can be represented as an array). <br>\nWhy no?<br>\nIf the notebook output contains more than 100,000 files - it is packed into a zip file.<br>\nOkay, we split it even smaller. Load 160 notebooks into the main code - we get an error that the notebook cannot be saved.<br>\n<strong>Second</strong><br>\nUsing your own zip files (ZIP64, tar, zip7). Folders with pkl files in zip.<br>\nWhy no?<br>\nOpening a ZIP file takes a few seconds and hundreds of megabytes of RAM.</p>\n<h2>Why it can be useful?</h2>\n<p>Most patents have a title and description rather than abstract or claims. Now you can get the description (the most complete information) of each patent for 0.1 ms from test.csv when you submit. I have my notebook running for an average of 5 hours.</p>\n<h2>Tip</h2>\n<p>There are 2 weeks left till the end of the competition, I hope this will help someone.</p>",
      "rawMarkdown": "## Problem\nOne of the problems with this contest is the use of competition data. If you want to extract data from a single patent, you have to access patent_data/year_month.parquet and read the data. Yes, you can use tricks like reading inline csv, using pyarrow, Rust libraries and other things. But the problem is the time and RAM of loading the information of a single patent (tens of seconds and gigabytes of memory).\n## Solution\nThe solution to the problem is to use HDF5 files to store patent information. Namely, storing patent information for n-number of years in multiple files (e.g. 2000_1.h5) and files in a notebook output (up to 20 GB). Loading these notebooks into your main code. HDF5 provide access to data as in hashmap for O(1) time, in my case by 'publication_nubmer'. Open the file - get the information of the desired patent without loading the rest of the patents.\n## Links to finished notebooks (11 in total)\nEach notebook contains patents from year **i up j**\nhttps://www.kaggle.com/code/qurusx/1700up1929\nhttps://www.kaggle.com/code/qurusx/1930up1974\nhttps://www.kaggle.com/code/qurusx/1975up1994\nhttps://www.kaggle.com/code/qurusx/1995up2004\nhttps://www.kaggle.com/code/qurusx/2005up2009\nhttps://www.kaggle.com/code/qurusx/2010up2013\nhttps://www.kaggle.com/code/qurusx/2014up2016\nhttps://www.kaggle.com/code/qurusx/2017up2019\nhttps://www.kaggle.com/code/qurusx/2020up2021\nhttps://www.kaggle.com/code/qurusx/2022up2022\nhttps://www.kaggle.com/code/qurusx/2023up2023\n## A link to a notebook with an example of how to use it\nhttps://www.kaggle.com/code/qurusx/description-fast-extract\n## What ideas didn't work (if interested)\n**The first one**\nUsing pkl files. We form the parquet into a folder with the same name, and each line in it into a separate pkl file (a line can be represented as an array). \nWhy no?\nIf the notebook output contains more than 100,000 files - it is packed into a zip file.\nOkay, we split it even smaller. Load 160 notebooks into the main code - we get an error that the notebook cannot be saved.\n**Second**\nUsing your own zip files (ZIP64, tar, zip7). Folders with pkl files in zip.\nWhy no?\nOpening a ZIP file takes a few seconds and hundreds of megabytes of RAM.\n## Why it can be useful?\nMost patents have a title and description rather than abstract or claims. Now you can get the description (the most complete information) of each patent for 0.1 ms from test.csv when you submit. I have my notebook running for an average of 5 hours.\n## Tip\nThere are 2 weeks left till the end of the competition, I hope this will help someone.",
      "votes": null
    },
    {
      "id": "2916432",
      "postDate": "07/11/2024 02:56:03",
      "content": "<p>This will be really helpful! Thanks so much <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> ☺️</p>",
      "rawMarkdown": "This will be really helpful! Thanks so much @qurusx ☺️",
      "votes": null
    },
    {
      "id": "2918237",
      "postDate": "07/12/2024 05:26:57",
      "content": "<p>Thank you for publishing so useful notebook!<br>\nI have been struggling to extract description for weeks:) (and I can't do it for now) </p>",
      "rawMarkdown": "Thank you for publishing so useful notebook!\nI have been struggling to extract description for weeks:) (and I can't do it for now)",
      "votes": null
    },
    {
      "id": "2918858",
      "postDate": "07/12/2024 14:52:38",
      "content": "<p>Nice trick with the HDF5 files! I don't know if it's me, but the links to the notebooks seem to be broken.</p>",
      "rawMarkdown": "Nice trick with the HDF5 files! I don't know if it's me, but the links to the notebooks seem to be broken.",
      "votes": null
    },
    {
      "id": "2919065",
      "postDate": "07/12/2024 16:57:29",
      "content": "<p>Thanks for the note. Also thanks for clearing RAM using ctypes - that helped with loading a few large parquets. Everything should be working now!</p>",
      "rawMarkdown": "Thanks for the note. Also thanks for clearing RAM using ctypes - that helped with loading a few large parquets. Everything should be working now!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2916432,
      "author_name": "hopefulleee",
      "author_url": "",
      "post_date": "07/11/2024 02:56:03",
      "content": "<p>This will be really helpful! Thanks so much <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> ☺️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2918237,
      "author_name": "aioraito",
      "author_url": "",
      "post_date": "07/12/2024 05:26:57",
      "content": "<p>Thank you for publishing so useful notebook!<br>\nI have been struggling to extract description for weeks:) (and I can't do it for now) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2918858,
      "author_name": "apparition",
      "author_url": "",
      "post_date": "07/12/2024 14:52:38",
      "content": "<p>Nice trick with the HDF5 files! I don't know if it's me, but the links to the notebooks seem to be broken.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2919065,
          "author_name": "qurusx",
          "author_url": "",
          "post_date": "07/12/2024 16:57:29",
          "content": "<p>Thanks for the note. Also thanks for clearing RAM using ctypes - that helped with loading a few large parquets. Everything should be working now!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2916129": "## Problem\nOne of the problems with this contest is the use of competition data. If you want to extract data from a single patent, you have to access patent_data/year_month.parquet and read the data. Yes, you can use tricks like reading inline csv, using pyarrow, Rust libraries and other things. But the problem is the time and RAM of loading the information of a single patent (tens of seconds and gigabytes of memory).\n## Solution\nThe solution to the problem is to use HDF5 files to store patent information. Namely, storing patent information for n-number of years in multiple files (e.g. 2000_1.h5) and files in a notebook output (up to 20 GB). Loading these notebooks into your main code. HDF5 provide access to data as in hashmap for O(1) time, in my case by 'publication_nubmer'. Open the file - get the information of the desired patent without loading the rest of the patents.\n## Links to finished notebooks (11 in total)\nEach notebook contains patents from year **i up j**\nhttps://www.kaggle.com/code/qurusx/1700up1929\nhttps://www.kaggle.com/code/qurusx/1930up1974\nhttps://www.kaggle.com/code/qurusx/1975up1994\nhttps://www.kaggle.com/code/qurusx/1995up2004\nhttps://www.kaggle.com/code/qurusx/2005up2009\nhttps://www.kaggle.com/code/qurusx/2010up2013\nhttps://www.kaggle.com/code/qurusx/2014up2016\nhttps://www.kaggle.com/code/qurusx/2017up2019\nhttps://www.kaggle.com/code/qurusx/2020up2021\nhttps://www.kaggle.com/code/qurusx/2022up2022\nhttps://www.kaggle.com/code/qurusx/2023up2023\n## A link to a notebook with an example of how to use it\nhttps://www.kaggle.com/code/qurusx/description-fast-extract\n## What ideas didn't work (if interested)\n**The first one**\nUsing pkl files. We form the parquet into a folder with the same name, and each line in it into a separate pkl file (a line can be represented as an array). \nWhy no?\nIf the notebook output contains more than 100,000 files - it is packed into a zip file.\nOkay, we split it even smaller. Load 160 notebooks into the main code - we get an error that the notebook cannot be saved.\n**Second**\nUsing your own zip files (ZIP64, tar, zip7). Folders with pkl files in zip.\nWhy no?\nOpening a ZIP file takes a few seconds and hundreds of megabytes of RAM.\n## Why it can be useful?\nMost patents have a title and description rather than abstract or claims. Now you can get the description (the most complete information) of each patent for 0.1 ms from test.csv when you submit. I have my notebook running for an average of 5 hours.\n## Tip\nThere are 2 weeks left till the end of the competition, I hope this will help someone.",
    "2916432": "This will be really helpful! Thanks so much @qurusx ☺️",
    "2918237": "Thank you for publishing so useful notebook!\nI have been struggling to extract description for weeks:) (and I can't do it for now)",
    "2918858": "Nice trick with the HDF5 files! I don't know if it's me, but the links to the notebooks seem to be broken.",
    "2919065": "Thanks for the note. Also thanks for clearing RAM using ctypes - that helped with loading a few large parquets. Everything should be working now!"
  },
  "source": "meta"
}