{
  "id": 501381,
  "title": "duckdb vs. other options - reasons to use duckdb ",
  "url": "/competitions/leash-BELKA/discussion/501381",
  "author_name": "",
  "post_date": "2024-05-09T01:32:08.410216400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Why most kagglers are using duckdb whereas other options are available. For example, using following code, we can read specific number of records and specific number of features/ columns at a time. </p>\n<p>import pandas as pd</p>\n<h1>Define the columns to read from the CSV file</h1>\n<p>cols_to_read = ['molecule_smiles', 'protein_name', 'binds']<br>\nchunksize = 5  # may be adjusted as needed<br>\nchunk_no=1<br>\nchunks_to_read=2 # may be adjusted as needed </p>\n<h1>Read the CSV file in chunks</h1>\n<p>for chunk in pd.read_csv('train.csv', usecols=cols_to_read, chunksize=chunksize):</p>\n<pre><code>\n\n\n\nchunk_no+=\n\n chunks_to_read == :\n    break\nchunks_to_read-=\n</code></pre>",
  "messages": [
    {
      "id": "2802297",
      "postDate": "05/09/2024 01:32:08",
      "content": "<p>Why most kagglers are using duckdb whereas other options are available. For example, using following code, we can read specific number of records and specific number of features/ columns at a time. </p>\n<p>import pandas as pd</p>\n<h1>Define the columns to read from the CSV file</h1>\n<p>cols_to_read = ['molecule_smiles', 'protein_name', 'binds']<br>\nchunksize = 5  # may be adjusted as needed<br>\nchunk_no=1<br>\nchunks_to_read=2 # may be adjusted as needed </p>\n<h1>Read the CSV file in chunks</h1>\n<p>for chunk in pd.read_csv('train.csv', usecols=cols_to_read, chunksize=chunksize):</p>\n<pre><code>\n\n\n\nchunk_no+=\n\n chunks_to_read == :\n    break\nchunks_to_read-=\n</code></pre>",
      "rawMarkdown": "Why most kagglers are using duckdb whereas other options are available. For example, using following code, we can read specific number of records and specific number of features/ columns at a time. \n\nimport pandas as pd\n\n# Define the columns to read from the CSV file\ncols_to_read = ['molecule_smiles', 'protein_name', 'binds']\nchunksize = 5  # may be adjusted as needed\nchunk_no=1\nchunks_to_read=2 # may be adjusted as needed \n\n# Read the CSV file in chunks\nfor chunk in pd.read_csv('train.csv', usecols=cols_to_read, chunksize=chunksize):\n    \n    print('chunk number: ', chunk_no)\n    print(chunk['molecule_smiles'])\n    print(chunk['protein_name'])\n    print(chunk['binds'])\n    chunk_no+=1\n    \n    if chunks_to_read == 0:\n        break\n    chunks_to_read-=1",
      "votes": null
    },
    {
      "id": "2802391",
      "postDate": "05/09/2024 02:53:36",
      "content": "<p>Probably it's just what gets used first that works. A lot of cloned or copy/pasted code. </p>",
      "rawMarkdown": "Probably it's just what gets used first that works. A lot of cloned or copy/pasted code.",
      "votes": null
    },
    {
      "id": "2802420",
      "postDate": "05/09/2024 03:33:19",
      "content": "<p>Yes. You are right. Thinking on different lines brings innovation actually</p>",
      "rawMarkdown": "Yes. You are right. Thinking on different lines brings innovation actually",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2802391,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "05/09/2024 02:53:36",
      "content": "<p>Probably it's just what gets used first that works. A lot of cloned or copy/pasted code. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2802420,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "05/09/2024 03:33:19",
          "content": "<p>Yes. You are right. Thinking on different lines brings innovation actually</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2802297": "Why most kagglers are using duckdb whereas other options are available. For example, using following code, we can read specific number of records and specific number of features/ columns at a time. \n\nimport pandas as pd\n\n# Define the columns to read from the CSV file\ncols_to_read = ['molecule_smiles', 'protein_name', 'binds']\nchunksize = 5  # may be adjusted as needed\nchunk_no=1\nchunks_to_read=2 # may be adjusted as needed \n\n# Read the CSV file in chunks\nfor chunk in pd.read_csv('train.csv', usecols=cols_to_read, chunksize=chunksize):\n    \n    print('chunk number: ', chunk_no)\n    print(chunk['molecule_smiles'])\n    print(chunk['protein_name'])\n    print(chunk['binds'])\n    chunk_no+=1\n    \n    if chunks_to_read == 0:\n        break\n    chunks_to_read-=1",
    "2802391": "Probably it's just what gets used first that works. A lot of cloned or copy/pasted code.",
    "2802420": "Yes. You are right. Thinking on different lines brings innovation actually"
  },
  "source": "meta"
}