{
  "id": 497924,
  "title": "test.csv should contain more information",
  "url": "/competitions/uspto-explainable-ai/discussion/497924",
  "author_name": "",
  "post_date": "2024-04-26T08:33:07.680095400Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Edit: Nvm I figured it out</p>\n<p>Right now, the way I process <code>test.csv</code> is:</p>\n<ul>\n<li>Load <code>test.csv</code></li>\n<li>Load <code>patent_metadata.parquet</code></li>\n</ul>\n<p>For each row of <code>test.csv</code>:</p>\n<ul>\n<li>Cross check the publication date from metadata.</li>\n<li>Load the corresponding <code>patent_data/{year}_{month}.parquet</code> file.</li>\n<li>Pick the row with the right publication number to get the corresponding text.</li>\n</ul>\n<p>Please correct me if I am doing something wrong here, put this feels like a pointless and time-wasting procedure when the information from the <code>patent_data</code> files can just be embedded into <code>test.csv</code> instead.</p>",
  "messages": [
    {
      "id": "2776574",
      "postDate": "04/26/2024 08:33:07",
      "content": "<p>Edit: Nvm I figured it out</p>\n<p>Right now, the way I process <code>test.csv</code> is:</p>\n<ul>\n<li>Load <code>test.csv</code></li>\n<li>Load <code>patent_metadata.parquet</code></li>\n</ul>\n<p>For each row of <code>test.csv</code>:</p>\n<ul>\n<li>Cross check the publication date from metadata.</li>\n<li>Load the corresponding <code>patent_data/{year}_{month}.parquet</code> file.</li>\n<li>Pick the row with the right publication number to get the corresponding text.</li>\n</ul>\n<p>Please correct me if I am doing something wrong here, put this feels like a pointless and time-wasting procedure when the information from the <code>patent_data</code> files can just be embedded into <code>test.csv</code> instead.</p>",
      "rawMarkdown": "Edit: Nvm I figured it out\n\nRight now, the way I process `test.csv` is:\n- Load `test.csv`\n- Load `patent_metadata.parquet`\n\nFor each row of `test.csv`:\n- Cross check the publication date from metadata.\n- Load the corresponding `patent_data/{year}_{month}.parquet` file.\n- Pick the row with the right publication number to get the corresponding text.\n\nPlease correct me if I am doing something wrong here, put this feels like a pointless and time-wasting procedure when the information from the `patent_data` files can just be embedded into `test.csv` instead.",
      "votes": null
    },
    {
      "id": "2782692",
      "postDate": "04/29/2024 12:21:30",
      "content": "<p>oh that's exactly what I am doing!! But the problem is that even after deleting the loaded parquet patent_data/{year_date} file, mu CPU ram doesn't get free. how to clear CPU ram?</p>",
      "rawMarkdown": "oh that's exactly what I am doing!! But the problem is that even after deleting the loaded parquet patent_data/{year_date} file, mu CPU ram doesn't get free. how to clear CPU ram?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2782692,
      "author_name": "hunainimran",
      "author_url": "",
      "post_date": "04/29/2024 12:21:30",
      "content": "<p>oh that's exactly what I am doing!! But the problem is that even after deleting the loaded parquet patent_data/{year_date} file, mu CPU ram doesn't get free. how to clear CPU ram?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2776574": "Edit: Nvm I figured it out\n\nRight now, the way I process `test.csv` is:\n- Load `test.csv`\n- Load `patent_metadata.parquet`\n\nFor each row of `test.csv`:\n- Cross check the publication date from metadata.\n- Load the corresponding `patent_data/{year}_{month}.parquet` file.\n- Pick the row with the right publication number to get the corresponding text.\n\nPlease correct me if I am doing something wrong here, put this feels like a pointless and time-wasting procedure when the information from the `patent_data` files can just be embedded into `test.csv` instead.",
    "2782692": "oh that's exactly what I am doing!! But the problem is that even after deleting the loaded parquet patent_data/{year_date} file, mu CPU ram doesn't get free. how to clear CPU ram?"
  },
  "source": "meta"
}