{
  "id": 499639,
  "title": "Dataset to efficiently read Nearest Neighbors",
  "url": "/competitions/uspto-explainable-ai/discussion/499639",
  "author_name": "",
  "post_date": "2024-05-02T15:29:35.416302900Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have converted the nearest neighbors csv file into <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-nearest-neighbors-matrix-compressed\" target=\"_blank\">numpy matrix</a>. Feel free to download and use it. It will be loaded instantly keeping much less space compared to the original file which takes long to load and consume a lot of memory.</p>\n<pre><code> numpy  np\n pandas  pd\n tqdm  tqdm\n\nN = \nchunk_size = \n\n\nall_pubs = pd.read_parquet()[].values\npub_dict = {name: i  i, name  (all_pubs)}\n\nnn_matrix = np.zeros(((all_pubs), N), dtype=np.int32)\n\n chunk_df  tqdm(pd.read_csv(, chunksize=chunk_size)):\n     chunk_array  chunk_df.values:\n        i = pub_dict[chunk_array[]]\n         j  (N):\n            nn_matrix[i, j] = pub_dict[chunk_array[ + j]]\n\nnp.save(, nn_matrix)\n</code></pre>\n<p>The above is how the dataset is created (It took around 8 minutes to create in my machine). The index (row number) in metadata represents the patents and nearest neighbors can be represented in an integer matrix. This is much more efficient to read in terms of memory and time.</p>\n<p>Note: Some of them don't have neighbors. You can check it via:<br>\n<code>((nn_matrix == 0).sum(axis=1) &gt; 1)</code></p>",
  "messages": [
    {
      "id": "2789186",
      "postDate": "05/02/2024 15:29:35",
      "content": "<p>I have converted the nearest neighbors csv file into <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-nearest-neighbors-matrix-compressed\" target=\"_blank\">numpy matrix</a>. Feel free to download and use it. It will be loaded instantly keeping much less space compared to the original file which takes long to load and consume a lot of memory.</p>\n<pre><code> numpy  np\n pandas  pd\n tqdm  tqdm\n\nN = \nchunk_size = \n\n\nall_pubs = pd.read_parquet()[].values\npub_dict = {name: i  i, name  (all_pubs)}\n\nnn_matrix = np.zeros(((all_pubs), N), dtype=np.int32)\n\n chunk_df  tqdm(pd.read_csv(, chunksize=chunk_size)):\n     chunk_array  chunk_df.values:\n        i = pub_dict[chunk_array[]]\n         j  (N):\n            nn_matrix[i, j] = pub_dict[chunk_array[ + j]]\n\nnp.save(, nn_matrix)\n</code></pre>\n<p>The above is how the dataset is created (It took around 8 minutes to create in my machine). The index (row number) in metadata represents the patents and nearest neighbors can be represented in an integer matrix. This is much more efficient to read in terms of memory and time.</p>\n<p>Note: Some of them don't have neighbors. You can check it via:<br>\n<code>((nn_matrix == 0).sum(axis=1) &gt; 1)</code></p>",
      "rawMarkdown": "I have converted the nearest neighbors csv file into [numpy matrix](https://www.kaggle.com/datasets/aerdem4/uspto-nearest-neighbors-matrix-compressed). Feel free to download and use it. It will be loaded instantly keeping much less space compared to the original file which takes long to load and consume a lot of memory.\n\n```python\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\nN = 50\nchunk_size = 1000\n\n\nall_pubs = pd.read_parquet(\"data/patent_metadata.parquet\")[\"publication_number\"].values\npub_dict = {name: i for i, name in enumerate(all_pubs)}\n\nnn_matrix = np.zeros((len(all_pubs), N), dtype=np.int32)\n\nfor chunk_df in tqdm(pd.read_csv(\"data/nearest_neighbors.csv\", chunksize=chunk_size)):\n    for chunk_array in chunk_df.values:\n        i = pub_dict[chunk_array[0]]\n        for j in range(N):\n            nn_matrix[i, j] = pub_dict[chunk_array[1 + j]]\n\nnp.save(\"cache/nearest_neighbors_matrix.npy\", nn_matrix)\n```\n\nThe above is how the dataset is created (It took around 8 minutes to create in my machine). The index (row number) in metadata represents the patents and nearest neighbors can be represented in an integer matrix. This is much more efficient to read in terms of memory and time.\n\nNote: Some of them don't have neighbors. You can check it via:\n`((nn_matrix == 0).sum(axis=1) &gt; 1)`",
      "votes": null
    },
    {
      "id": "2789248",
      "postDate": "05/02/2024 16:09:28",
      "content": "<p>Nice to see publication numbers are replaced with its index, Thanks for sharing <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> </p>",
      "rawMarkdown": "Nice to see publication numbers are replaced with its index, Thanks for sharing @aerdem4",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2789248,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "05/02/2024 16:09:28",
      "content": "<p>Nice to see publication numbers are replaced with its index, Thanks for sharing <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2789186": "I have converted the nearest neighbors csv file into [numpy matrix](https://www.kaggle.com/datasets/aerdem4/uspto-nearest-neighbors-matrix-compressed). Feel free to download and use it. It will be loaded instantly keeping much less space compared to the original file which takes long to load and consume a lot of memory.\n\n```python\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\nN = 50\nchunk_size = 1000\n\n\nall_pubs = pd.read_parquet(\"data/patent_metadata.parquet\")[\"publication_number\"].values\npub_dict = {name: i for i, name in enumerate(all_pubs)}\n\nnn_matrix = np.zeros((len(all_pubs), N), dtype=np.int32)\n\nfor chunk_df in tqdm(pd.read_csv(\"data/nearest_neighbors.csv\", chunksize=chunk_size)):\n    for chunk_array in chunk_df.values:\n        i = pub_dict[chunk_array[0]]\n        for j in range(N):\n            nn_matrix[i, j] = pub_dict[chunk_array[1 + j]]\n\nnp.save(\"cache/nearest_neighbors_matrix.npy\", nn_matrix)\n```\n\nThe above is how the dataset is created (It took around 8 minutes to create in my machine). The index (row number) in metadata represents the patents and nearest neighbors can be represented in an integer matrix. This is much more efficient to read in terms of memory and time.\n\nNote: Some of them don't have neighbors. You can check it via:\n`((nn_matrix == 0).sum(axis=1) &gt; 1)`",
    "2789248": "Nice to see publication numbers are replaced with its index, Thanks for sharing @aerdem4"
  },
  "source": "meta"
}