{
  "id": 68541,
  "title": "Test Set csv in sqlite format",
  "url": "/competitions/PLAsTiCC-2018/discussion/68541",
  "author_name": "",
  "post_date": "2018-10-14T08:52:36.135090600Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I was just wondering it is was possible to have test_set.csv in sqlite format.</p>\n\n<p>To make things fit in a kernel like I published <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\">here</a> the main issue is to read test samples at the object_id level and as far as I known pandas does not provide such a functionality.</p>\n\n<p>To do this from a csv file one would need to read the file line by line... which I will end up doing.</p>\n\n<p>I believe sqlite format would ease that a bit since we would be able to 1st read the object_ids and then run a few selects to get groups of objects. Though that may be time consuming...</p>\n\n<p>Don't know if sharing a sqlite file is feasible in the current kernel framework but I think that would be interesting.</p>\n\n<p>What are your thoughts ? </p>",
  "messages": [
    {
      "id": "403656",
      "postDate": "10/14/2018 08:52:36",
      "content": "<p>I was just wondering it is was possible to have test_set.csv in sqlite format.</p>\n\n<p>To make things fit in a kernel like I published <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\">here</a> the main issue is to read test samples at the object_id level and as far as I known pandas does not provide such a functionality.</p>\n\n<p>To do this from a csv file one would need to read the file line by line... which I will end up doing.</p>\n\n<p>I believe sqlite format would ease that a bit since we would be able to 1st read the object_ids and then run a few selects to get groups of objects. Though that may be time consuming...</p>\n\n<p>Don't know if sharing a sqlite file is feasible in the current kernel framework but I think that would be interesting.</p>\n\n<p>What are your thoughts ? </p>",
      "rawMarkdown": "I was just wondering it is was possible to have test_set.csv in sqlite format.\n\nTo make things fit in a kernel like I published [here](https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data) the main issue is to read test samples at the object_id level and as far as I known pandas does not provide such a functionality.\n\nTo do this from a csv file one would need to read the file line by line... which I will end up doing.\n\nI believe sqlite format would ease that a bit since we would be able to 1st read the object_ids and then run a few selects to get groups of objects. Though that may be time consuming...\n\nDon't know if sharing a sqlite file is feasible in the current kernel framework but I think that would be interesting.\n\nWhat are your thoughts ?",
      "votes": null
    },
    {
      "id": "403658",
      "postDate": "10/14/2018 09:00:20",
      "content": "<p>Hey Olivier,</p>\n\n<p>I'm pretty sure that doing a <code>df.groupby</code> on a <code>pandas.DataFrame</code> read in chunks will produce the correct results. I guess <code>pandas</code> handles this intelligently. See <a href=\"https://stackoverflow.com/questions/23190156/pandas-groupby-mean-of-large-dataset-in-csv\">here</a> for an example. </p>",
      "rawMarkdown": "Hey Olivier,\n\nI'm pretty sure that doing a `df.groupby` on a `pandas.DataFrame` read in chunks will produce the correct results. I guess `pandas` handles this intelligently. See [here](https://stackoverflow.com/questions/23190156/pandas-groupby-mean-of-large-dataset-in-csv) for an example.",
      "votes": null
    },
    {
      "id": "403672",
      "postDate": "10/14/2018 09:47:02",
      "content": "<p>I don't think it prevents having object_ids spanning 2 consecutive chunks.</p>\n\n<p>My kernel ends up with 88 duplicate object_ids doing the groupby in chunks.</p>",
      "rawMarkdown": "I don't think it prevents having object_ids spanning 2 consecutive chunks.\n\nMy kernel ends up with 88 duplicate object_ids doing the groupby in chunks.",
      "votes": null
    },
    {
      "id": "403673",
      "postDate": "10/14/2018 09:54:05",
      "content": "<p>My bad :). Good to know though.</p>",
      "rawMarkdown": "My bad :). Good to know though.",
      "votes": null
    },
    {
      "id": "403676",
      "postDate": "10/14/2018 09:59:31",
      "content": "<p>The solution you link to works I'm just not sure the whole thing would fit in memory. I already had issues with the same script in a notebook...</p>",
      "rawMarkdown": "The solution you link to works I'm just not sure the whole thing would fit in memory. I already had issues with the same script in a notebook...",
      "votes": null
    },
    {
      "id": "403844",
      "postDate": "10/14/2018 17:53:43",
      "content": "<p>Hey Olivier,</p>\n\n<p>what about checking if the last object_id of a chunk is part of the next chunk? Pandas read_csv has a usecols parameter, which is useful if you are only interested in particular columns. With this approach one would learn where to make the splits. </p>\n\n<p>If this is not what you are looking for, maybe it is worth, to have a look at the python dask library, which focusses on handling large data sets and parallelization.</p>",
      "rawMarkdown": "Hey Olivier,\n\nwhat about checking if the last object_id of a chunk is part of the next chunk? Pandas read_csv has a usecols parameter, which is useful if you are only interested in particular columns. With this approach one would learn where to make the splits. \n\nIf this is not what you are looking for, maybe it is worth, to have a look at the python dask library, which focusses on handling large data sets and parallelization.",
      "votes": null
    },
    {
      "id": "403870",
      "postDate": "10/14/2018 19:17:03",
      "content": "<p>It is a good idea to have test_set.csv in sqlite format!</p>",
      "rawMarkdown": "It is a good idea to have test_set.csv in sqlite format!",
      "votes": null
    },
    {
      "id": "405529",
      "postDate": "10/17/2018 16:40:11",
      "content": "<p>Hi Olivier,</p>\n\n<p>I saved the test_set.csv data to an sqlite database. I hoped to use pandas.read_sql with a \"GROUP BY\" in the query, but that didn't work because (it appears) pandas wants to read in the entire data set, thus too large to fit my memory.</p>\n\n<p>I've settled, for now, on looping over the list of object_ids, forming a sublist of N (e.g. 20,000) IDs, then submitting a pandas.read_sql query only on that sublist. From there, I use the pandas.DataFrame.groupby to get the correct data. It's slower than I'd like, but seems to work.</p>\n\n<p>I'm open to suggestions!</p>",
      "rawMarkdown": "Hi Olivier,\n\nI saved the test_set.csv data to an sqlite database. I hoped to use pandas.read_sql with a \"GROUP BY\" in the query, but that didn't work because (it appears) pandas wants to read in the entire data set, thus too large to fit my memory.\n\nI've settled, for now, on looping over the list of object_ids, forming a sublist of N (e.g. 20,000) IDs, then submitting a pandas.read_sql query only on that sublist. From there, I use the pandas.DataFrame.groupby to get the correct data. It's slower than I'd like, but seems to work.\n\nI'm open to suggestions!",
      "votes": null
    },
    {
      "id": "405549",
      "postDate": "10/17/2018 17:25:39",
      "content": "<p>There are 2 public kernels that sort this issue : </p>\n\n<p><a href=\"https://www.kaggle.com/johnfarrell/plasticc-in-a-kernel-meta-and-data-c7a0fc\">https://www.kaggle.com/johnfarrell/plasticc-in-a-kernel-meta-and-data-c7a0fc</a></p>\n\n<p>and mine <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data</a></p>\n\n<p>they use pandas read in chunk and a special process to avoid having objects shared on 2 chunks.</p>\n\n<p>They work fine as  far as I can tell (I'm in kernel only format for now)  </p>",
      "rawMarkdown": "There are 2 public kernels that sort this issue : \n\nhttps://www.kaggle.com/johnfarrell/plasticc-in-a-kernel-meta-and-data-c7a0fc\n\nand mine https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\n\nthey use pandas read in chunk and a special process to avoid having objects shared on 2 chunks.\n\nThey work fine as  far as I can tell (I'm in kernel only format for now)",
      "votes": null
    },
    {
      "id": "405572",
      "postDate": "10/17/2018 18:15:05",
      "content": "<p>Thank you, Olivier. I'll have a look and compare with what I've done.</p>",
      "rawMarkdown": "Thank you, Olivier. I'll have a look and compare with what I've done.",
      "votes": null
    },
    {
      "id": "407865",
      "postDate": "10/21/2018 22:49:37",
      "content": "<p>hi Olivier</p>\n\n<p>food for thoughts. not directly answering your SQLite question but I think it is related,</p>\n\n<p>I processed by set of object Id ending with the same digit,\nthis way I had sets of objects ids and their time series that fits easily in memory even on a small laptop,</p>\n\n<p>Another advantage is that you process the entire time series for each object in the set  ( not easy when you process by chunks on the test set directly),</p>\n\n<p>Jc </p>",
      "rawMarkdown": "hi Olivier\n\nfood for thoughts. not directly answering your SQLite question but I think it is related,\n\nI processed by set of object Id ending with the same digit,\nthis way I had sets of objects ids and their time series that fits easily in memory even on a small laptop,\n\nAnother advantage is that you process the entire time series for each object in the set  ( not easy when you process by chunks on the test set directly),\n\nJc",
      "votes": null
    },
    {
      "id": "408370",
      "postDate": "10/22/2018 18:27:58",
      "content": "<p>You may find <a href=\"https://www.kaggle.com/alexfir/fast-test-set-reading\">Fast test_set reading</a> kernel useful.</p>",
      "rawMarkdown": "You may find [Fast test_set reading][1] kernel useful.\n\n\n  [1]: https://www.kaggle.com/alexfir/fast-test-set-reading",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 403658,
      "author_name": "maxhalford",
      "author_url": "",
      "post_date": "10/14/2018 09:00:20",
      "content": "<p>Hey Olivier,</p>\n\n<p>I'm pretty sure that doing a <code>df.groupby</code> on a <code>pandas.DataFrame</code> read in chunks will produce the correct results. I guess <code>pandas</code> handles this intelligently. See <a href=\"https://stackoverflow.com/questions/23190156/pandas-groupby-mean-of-large-dataset-in-csv\">here</a> for an example. </p>",
      "votes": null,
      "replies": [
        {
          "id": 403672,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "10/14/2018 09:47:02",
          "content": "<p>I don't think it prevents having object_ids spanning 2 consecutive chunks.</p>\n\n<p>My kernel ends up with 88 duplicate object_ids doing the groupby in chunks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 403673,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "10/14/2018 09:54:05",
          "content": "<p>My bad :). Good to know though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 403676,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "10/14/2018 09:59:31",
          "content": "<p>The solution you link to works I'm just not sure the whole thing would fit in memory. I already had issues with the same script in a notebook...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405529,
          "author_name": "gregbehm",
          "author_url": "",
          "post_date": "10/17/2018 16:40:11",
          "content": "<p>Hi Olivier,</p>\n\n<p>I saved the test_set.csv data to an sqlite database. I hoped to use pandas.read_sql with a \"GROUP BY\" in the query, but that didn't work because (it appears) pandas wants to read in the entire data set, thus too large to fit my memory.</p>\n\n<p>I've settled, for now, on looping over the list of object_ids, forming a sublist of N (e.g. 20,000) IDs, then submitting a pandas.read_sql query only on that sublist. From there, I use the pandas.DataFrame.groupby to get the correct data. It's slower than I'd like, but seems to work.</p>\n\n<p>I'm open to suggestions!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405549,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "10/17/2018 17:25:39",
          "content": "<p>There are 2 public kernels that sort this issue : </p>\n\n<p><a href=\"https://www.kaggle.com/johnfarrell/plasticc-in-a-kernel-meta-and-data-c7a0fc\">https://www.kaggle.com/johnfarrell/plasticc-in-a-kernel-meta-and-data-c7a0fc</a></p>\n\n<p>and mine <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data</a></p>\n\n<p>they use pandas read in chunk and a special process to avoid having objects shared on 2 chunks.</p>\n\n<p>They work fine as  far as I can tell (I'm in kernel only format for now)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405572,
          "author_name": "gregbehm",
          "author_url": "",
          "post_date": "10/17/2018 18:15:05",
          "content": "<p>Thank you, Olivier. I'll have a look and compare with what I've done.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 403844,
      "author_name": "jeffkk",
      "author_url": "",
      "post_date": "10/14/2018 17:53:43",
      "content": "<p>Hey Olivier,</p>\n\n<p>what about checking if the last object_id of a chunk is part of the next chunk? Pandas read_csv has a usecols parameter, which is useful if you are only interested in particular columns. With this approach one would learn where to make the splits. </p>\n\n<p>If this is not what you are looking for, maybe it is worth, to have a look at the python dask library, which focusses on handling large data sets and parallelization.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 403870,
      "author_name": "darbin",
      "author_url": "",
      "post_date": "10/14/2018 19:17:03",
      "content": "<p>It is a good idea to have test_set.csv in sqlite format!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 407865,
      "author_name": "jmeriaux",
      "author_url": "",
      "post_date": "10/21/2018 22:49:37",
      "content": "<p>hi Olivier</p>\n\n<p>food for thoughts. not directly answering your SQLite question but I think it is related,</p>\n\n<p>I processed by set of object Id ending with the same digit,\nthis way I had sets of objects ids and their time series that fits easily in memory even on a small laptop,</p>\n\n<p>Another advantage is that you process the entire time series for each object in the set  ( not easy when you process by chunks on the test set directly),</p>\n\n<p>Jc </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 408370,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "10/22/2018 18:27:58",
      "content": "<p>You may find <a href=\"https://www.kaggle.com/alexfir/fast-test-set-reading\">Fast test_set reading</a> kernel useful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "403656": "I was just wondering it is was possible to have test_set.csv in sqlite format.\n\nTo make things fit in a kernel like I published [here](https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data) the main issue is to read test samples at the object_id level and as far as I known pandas does not provide such a functionality.\n\nTo do this from a csv file one would need to read the file line by line... which I will end up doing.\n\nI believe sqlite format would ease that a bit since we would be able to 1st read the object_ids and then run a few selects to get groups of objects. Though that may be time consuming...\n\nDon't know if sharing a sqlite file is feasible in the current kernel framework but I think that would be interesting.\n\nWhat are your thoughts ?",
    "403658": "Hey Olivier,\n\nI'm pretty sure that doing a `df.groupby` on a `pandas.DataFrame` read in chunks will produce the correct results. I guess `pandas` handles this intelligently. See [here](https://stackoverflow.com/questions/23190156/pandas-groupby-mean-of-large-dataset-in-csv) for an example.",
    "403672": "I don't think it prevents having object_ids spanning 2 consecutive chunks.\n\nMy kernel ends up with 88 duplicate object_ids doing the groupby in chunks.",
    "403673": "My bad :). Good to know though.",
    "403676": "The solution you link to works I'm just not sure the whole thing would fit in memory. I already had issues with the same script in a notebook...",
    "403844": "Hey Olivier,\n\nwhat about checking if the last object_id of a chunk is part of the next chunk? Pandas read_csv has a usecols parameter, which is useful if you are only interested in particular columns. With this approach one would learn where to make the splits. \n\nIf this is not what you are looking for, maybe it is worth, to have a look at the python dask library, which focusses on handling large data sets and parallelization.",
    "403870": "It is a good idea to have test_set.csv in sqlite format!",
    "405529": "Hi Olivier,\n\nI saved the test_set.csv data to an sqlite database. I hoped to use pandas.read_sql with a \"GROUP BY\" in the query, but that didn't work because (it appears) pandas wants to read in the entire data set, thus too large to fit my memory.\n\nI've settled, for now, on looping over the list of object_ids, forming a sublist of N (e.g. 20,000) IDs, then submitting a pandas.read_sql query only on that sublist. From there, I use the pandas.DataFrame.groupby to get the correct data. It's slower than I'd like, but seems to work.\n\nI'm open to suggestions!",
    "405549": "There are 2 public kernels that sort this issue : \n\nhttps://www.kaggle.com/johnfarrell/plasticc-in-a-kernel-meta-and-data-c7a0fc\n\nand mine https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\n\nthey use pandas read in chunk and a special process to avoid having objects shared on 2 chunks.\n\nThey work fine as  far as I can tell (I'm in kernel only format for now)",
    "405572": "Thank you, Olivier. I'll have a look and compare with what I've done.",
    "407865": "hi Olivier\n\nfood for thoughts. not directly answering your SQLite question but I think it is related,\n\nI processed by set of object Id ending with the same digit,\nthis way I had sets of objects ids and their time series that fits easily in memory even on a small laptop,\n\nAnother advantage is that you process the entire time series for each object in the set  ( not easy when you process by chunks on the test set directly),\n\nJc",
    "408370": "You may find [Fast test_set reading][1] kernel useful.\n\n\n  [1]: https://www.kaggle.com/alexfir/fast-test-set-reading"
  },
  "source": "meta"
}