{
  "id": 195078,
  "title": "Datatable Help",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195078",
  "author_name": "",
  "post_date": "2020-11-03T13:27:09.217146600Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I'm going to repeat some of the discussion from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377\" target=\"_blank\">How much time should a successful submission take?\n</a> as I think it would be useful to have all the questions and answers about datatable in one place. As I have more questions! 😊</p>\n<p><strong>Question 1</strong><br>\n<a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> do you know of any good tutorials for datatable? I've seen your <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">notebook on reading data in</a> and I'm reading the <a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">user guide/API</a>.<br>\nIt seem very quick, but I keep stumbling into issues.</p>\n<p>For example, how do I decrement the values in a column by 1?<br>\nI started by trying this</p>\n<pre><code>train['part'] = train['part'] - 1 \n</code></pre>\n<p>but got the following error</p>\n<pre><code>TypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n</code></pre>\n<p>This is what I have came up with, but it feels like I'm missing something.</p>\n<pre><code># Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n</code></pre>\n<p>On my local machine this takes 3ms vs 31ms in pandas (<code>train['part'] = train['part'] - 1</code>), so still alot quicker.</p>\n<p><strong>Answer from <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a></strong><br>\nThe datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.</p>\n<p>Till then you could use this forum or the <a href=\"https://github.com/h2oai/datatable/issues\" target=\"_blank\">Github repo</a>, I'm sure folks will be happy to help, me included.</p>\n<p><code>train['part'] = train[:, dt.f.part - 1]</code></p>\n<p>The <a href=\"https://datatable.readthedocs.io/en/latest/manual/f-expressions.html\" target=\"_blank\">f-expression</a> is used for most of these kind of operations.</p>\n<p><strong>Question 2</strong><br>\nThe documentation says it is missing the equivalent function to <code>cumsum</code> in pandas. Is there a better way of getting around than this as the conversion to/from pandas is quite slow.</p>\n<pre><code>train = train.to_pandas()\ntrain['Z'] = train.groupby(['X', 'Y']).cumcount()\ntrain = dt.Frame(train)\n</code></pre>\n<p>I'd be grateful for any help. Thanks</p>",
  "messages": [
    {
      "id": "1068497",
      "postDate": "11/03/2020 13:27:09",
      "content": "<p>Hello,</p>\n<p>I'm going to repeat some of the discussion from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377\" target=\"_blank\">How much time should a successful submission take?\n</a> as I think it would be useful to have all the questions and answers about datatable in one place. As I have more questions! 😊</p>\n<p><strong>Question 1</strong><br>\n<a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> do you know of any good tutorials for datatable? I've seen your <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">notebook on reading data in</a> and I'm reading the <a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">user guide/API</a>.<br>\nIt seem very quick, but I keep stumbling into issues.</p>\n<p>For example, how do I decrement the values in a column by 1?<br>\nI started by trying this</p>\n<pre><code>train['part'] = train['part'] - 1 \n</code></pre>\n<p>but got the following error</p>\n<pre><code>TypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n</code></pre>\n<p>This is what I have came up with, but it feels like I'm missing something.</p>\n<pre><code># Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n</code></pre>\n<p>On my local machine this takes 3ms vs 31ms in pandas (<code>train['part'] = train['part'] - 1</code>), so still alot quicker.</p>\n<p><strong>Answer from <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a></strong><br>\nThe datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.</p>\n<p>Till then you could use this forum or the <a href=\"https://github.com/h2oai/datatable/issues\" target=\"_blank\">Github repo</a>, I'm sure folks will be happy to help, me included.</p>\n<p><code>train['part'] = train[:, dt.f.part - 1]</code></p>\n<p>The <a href=\"https://datatable.readthedocs.io/en/latest/manual/f-expressions.html\" target=\"_blank\">f-expression</a> is used for most of these kind of operations.</p>\n<p><strong>Question 2</strong><br>\nThe documentation says it is missing the equivalent function to <code>cumsum</code> in pandas. Is there a better way of getting around than this as the conversion to/from pandas is quite slow.</p>\n<pre><code>train = train.to_pandas()\ntrain['Z'] = train.groupby(['X', 'Y']).cumcount()\ntrain = dt.Frame(train)\n</code></pre>\n<p>I'd be grateful for any help. Thanks</p>",
      "rawMarkdown": "Hello,\n\nI'm going to repeat some of the discussion from [How much time should a successful submission take?\n](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377) as I think it would be useful to have all the questions and answers about datatable in one place. As I have more questions! 😊\n\n**Question 1**\n@rohanrao do you know of any good tutorials for datatable? I've seen your [notebook on reading data in](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid) and I'm reading the [user guide/API](https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html).\nIt seem very quick, but I keep stumbling into issues.\n\nFor example, how do I decrement the values in a column by 1?\nI started by trying this\n```\ntrain['part'] = train['part'] - 1 \n```\nbut got the following error\n```\nTypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n```\nThis is what I have came up with, but it feels like I'm missing something.\n```\n# Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n```\nOn my local machine this takes 3ms vs 31ms in pandas (`train['part'] = train['part'] - 1`), so still alot quicker.\n\n**Answer from @rohanrao**\nThe datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.\n\nTill then you could use this forum or the [Github repo](https://github.com/h2oai/datatable/issues), I'm sure folks will be happy to help, me included.\n\n`train['part'] = train[:, dt.f.part - 1]`\n\nThe [f-expression](https://datatable.readthedocs.io/en/latest/manual/f-expressions.html) is used for most of these kind of operations.\n\n**Question 2**\nThe documentation says it is missing the equivalent function to `cumsum` in pandas. Is there a better way of getting around than this as the conversion to/from pandas is quite slow.\n```\ntrain = train.to_pandas()\ntrain['Z'] = train.groupby(['X', 'Y']).cumcount()\ntrain = dt.Frame(train)\n```\n\nI'd be grateful for any help. Thanks",
      "votes": null
    },
    {
      "id": "1068530",
      "postDate": "11/03/2020 13:42:47",
      "content": "<p>Cumulative functions aren't supported. But you can use pandas without having to convert the entire data. Just use the required columns. It shouldn't be very slow.</p>\n<p><code>train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()</code> (corrected from below)</p>",
      "rawMarkdown": "Cumulative functions aren't supported. But you can use pandas without having to convert the entire data. Just use the required columns. It shouldn't be very slow.\n\n`train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()` (corrected from below)",
      "votes": null
    },
    {
      "id": "1068704",
      "postDate": "11/03/2020 16:51:17",
      "content": "<blockquote>\n  <p>For example, how do I decrement the values in a column by 1?<br>\n  train['part'] = train['part'] - 1 </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/ghostskipper\" target=\"_blank\">@ghostskipper</a> Have you tried using an <code>f</code>-expression?:</p>\n<pre><code>train['part'] = dt.f.part - 1\n</code></pre>",
      "rawMarkdown": "> For example, how do I decrement the values in a column by 1?\n> train['part'] = train['part'] - 1 \n\n@ghostskipper Have you tried using an `f`-expression?:\n```\ntrain['part'] = dt.f.part - 1\n```",
      "votes": null
    },
    {
      "id": "1068717",
      "postDate": "11/03/2020 17:06:24",
      "content": "<p>Thank you! That's a bit neater and quicker.<br>\nMinor correction<br>\n<code>train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()</code></p>",
      "rawMarkdown": "Thank you! That's a bit neater and quicker.\nMinor correction\n`train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()`",
      "votes": null
    },
    {
      "id": "1069280",
      "postDate": "11/04/2020 09:28:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nyakaggle\" target=\"_blank\">@nyakaggle</a> I hadn't. Thanks for the suggestion, from my testing it appears to be faster.<br>\nFor the test I read the train.csv and added a dummy column (I added the dummy column before each test so they started from the same place)</p>\n<pre><code>train['part'] = 100001\ntrain[\"part\"] = dt.stype.int64(train[\"part\"])\n</code></pre>\n<p>For the first option</p>\n<pre><code>for ii in range(100000):\n    train['part'] = train[:, dt.f.part - 1]\n</code></pre>\n<p>This took 2m 12.76s</p>\n<p>The second option</p>\n<pre><code>for ii in range(100000):\n    train['part'] = dt.f.part - 1\n</code></pre>\n<p>This took 243ms</p>",
      "rawMarkdown": "Hi @nyakaggle I hadn't. Thanks for the suggestion, from my testing it appears to be faster.\nFor the test I read the train.csv and added a dummy column (I added the dummy column before each test so they started from the same place)\n```\ntrain['part'] = 100001\ntrain[\"part\"] = dt.stype.int64(train[\"part\"])\n```\nFor the first option\n```\nfor ii in range(100000):\n    train['part'] = train[:, dt.f.part - 1]\n```\nThis took 2m 12.76s\n\nThe second option\n```\nfor ii in range(100000):\n    train['part'] = dt.f.part - 1\n```\nThis took 243ms",
      "votes": null
    },
    {
      "id": "1069332",
      "postDate": "11/04/2020 10:18:29",
      "content": "<p>Ah yes this will be faster, thanks!</p>",
      "rawMarkdown": "Ah yes this will be faster, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1068530,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "11/03/2020 13:42:47",
      "content": "<p>Cumulative functions aren't supported. But you can use pandas without having to convert the entire data. Just use the required columns. It shouldn't be very slow.</p>\n<p><code>train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()</code> (corrected from below)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1068717,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/03/2020 17:06:24",
          "content": "<p>Thank you! That's a bit neater and quicker.<br>\nMinor correction<br>\n<code>train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1068704,
      "author_name": "nyakaggle",
      "author_url": "",
      "post_date": "11/03/2020 16:51:17",
      "content": "<blockquote>\n  <p>For example, how do I decrement the values in a column by 1?<br>\n  train['part'] = train['part'] - 1 </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/ghostskipper\" target=\"_blank\">@ghostskipper</a> Have you tried using an <code>f</code>-expression?:</p>\n<pre><code>train['part'] = dt.f.part - 1\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1069280,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/04/2020 09:28:29",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/nyakaggle\" target=\"_blank\">@nyakaggle</a> I hadn't. Thanks for the suggestion, from my testing it appears to be faster.<br>\nFor the test I read the train.csv and added a dummy column (I added the dummy column before each test so they started from the same place)</p>\n<pre><code>train['part'] = 100001\ntrain[\"part\"] = dt.stype.int64(train[\"part\"])\n</code></pre>\n<p>For the first option</p>\n<pre><code>for ii in range(100000):\n    train['part'] = train[:, dt.f.part - 1]\n</code></pre>\n<p>This took 2m 12.76s</p>\n<p>The second option</p>\n<pre><code>for ii in range(100000):\n    train['part'] = dt.f.part - 1\n</code></pre>\n<p>This took 243ms</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069332,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/04/2020 10:18:29",
          "content": "<p>Ah yes this will be faster, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1068497": "Hello,\n\nI'm going to repeat some of the discussion from [How much time should a successful submission take?\n](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377) as I think it would be useful to have all the questions and answers about datatable in one place. As I have more questions! 😊\n\n**Question 1**\n@rohanrao do you know of any good tutorials for datatable? I've seen your [notebook on reading data in](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid) and I'm reading the [user guide/API](https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html).\nIt seem very quick, but I keep stumbling into issues.\n\nFor example, how do I decrement the values in a column by 1?\nI started by trying this\n```\ntrain['part'] = train['part'] - 1 \n```\nbut got the following error\n```\nTypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n```\nThis is what I have came up with, but it feels like I'm missing something.\n```\n# Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n```\nOn my local machine this takes 3ms vs 31ms in pandas (`train['part'] = train['part'] - 1`), so still alot quicker.\n\n**Answer from @rohanrao**\nThe datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.\n\nTill then you could use this forum or the [Github repo](https://github.com/h2oai/datatable/issues), I'm sure folks will be happy to help, me included.\n\n`train['part'] = train[:, dt.f.part - 1]`\n\nThe [f-expression](https://datatable.readthedocs.io/en/latest/manual/f-expressions.html) is used for most of these kind of operations.\n\n**Question 2**\nThe documentation says it is missing the equivalent function to `cumsum` in pandas. Is there a better way of getting around than this as the conversion to/from pandas is quite slow.\n```\ntrain = train.to_pandas()\ntrain['Z'] = train.groupby(['X', 'Y']).cumcount()\ntrain = dt.Frame(train)\n```\n\nI'd be grateful for any help. Thanks",
    "1068530": "Cumulative functions aren't supported. But you can use pandas without having to convert the entire data. Just use the required columns. It shouldn't be very slow.\n\n`train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()` (corrected from below)",
    "1068704": "> For example, how do I decrement the values in a column by 1?\n> train['part'] = train['part'] - 1 \n\n@ghostskipper Have you tried using an `f`-expression?:\n```\ntrain['part'] = dt.f.part - 1\n```",
    "1068717": "Thank you! That's a bit neater and quicker.\nMinor correction\n`train['Z'] = train[:, ['X', 'Y']].to_pandas().groupby(['X', 'Y']).cumcount()`",
    "1069280": "Hi @nyakaggle I hadn't. Thanks for the suggestion, from my testing it appears to be faster.\nFor the test I read the train.csv and added a dummy column (I added the dummy column before each test so they started from the same place)\n```\ntrain['part'] = 100001\ntrain[\"part\"] = dt.stype.int64(train[\"part\"])\n```\nFor the first option\n```\nfor ii in range(100000):\n    train['part'] = train[:, dt.f.part - 1]\n```\nThis took 2m 12.76s\n\nThe second option\n```\nfor ii in range(100000):\n    train['part'] = dt.f.part - 1\n```\nThis took 243ms",
    "1069332": "Ah yes this will be faster, thanks!"
  },
  "source": "meta"
}