{
  "id": 194377,
  "title": "How much time should a successful submission take?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194377",
  "author_name": "",
  "post_date": "2020-11-01T14:32:47.342538800Z",
  "votes": 6,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Mine has been running since last 30mins.</p>",
  "messages": [
    {
      "id": "1066247",
      "postDate": "11/01/2020 14:32:47",
      "content": "<p>Mine has been running since last 30mins.</p>",
      "rawMarkdown": "Mine has been running since last 30mins.",
      "votes": null
    },
    {
      "id": "1066311",
      "postDate": "11/01/2020 15:47:09",
      "content": "<p>It will take time in hours. Data in test data set is huge</p>",
      "rawMarkdown": "It will take time in hours. Data in test data set is huge",
      "votes": null
    },
    {
      "id": "1066362",
      "postDate": "11/01/2020 17:03:35",
      "content": "<p>Not necessarily. It depends on features and structure of code (eg: I don't use pandas, it is super slow)<br>\nMine takes ~ 12 minutes for a successful submission.</p>",
      "rawMarkdown": "Not necessarily. It depends on features and structure of code (eg: I don't use pandas, it is super slow)\nMine takes ~ 12 minutes for a successful submission.",
      "votes": null
    },
    {
      "id": "1066373",
      "postDate": "11/01/2020 17:15:39",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> , Thanks for suggestion. As of now i am using pandas only and it is too slow 😏</p>",
      "rawMarkdown": "rohanrao , Thanks for suggestion. As of now i am using pandas only and it is too slow 😏",
      "votes": null
    },
    {
      "id": "1066391",
      "postDate": "11/01/2020 17:41:51",
      "content": "<p>I'm around 7hrs and I'm using pandas. Currently trying to move to something faster, I haven't figure out what to change to yet.</p>",
      "rawMarkdown": "I'm around 7hrs and I'm using pandas. Currently trying to move to something faster, I haven't figure out what to change to yet.",
      "votes": null
    },
    {
      "id": "1066683",
      "postDate": "11/01/2020 23:49:54",
      "content": "<blockquote>\n  <p>Mine has been running since last 30mins.</p>\n</blockquote>\n<p>Don't worry, I got successful submission after 2 to 3 hours of execution as well depending on models. 😊</p>",
      "rawMarkdown": "> Mine has been running since last 30mins.\n\nDon't worry, I got successful submission after 2 to 3 hours of execution as well depending on models. 😊",
      "votes": null
    },
    {
      "id": "1066688",
      "postDate": "11/01/2020 23:55:53",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> What would you recommend replacing pandas with to speed it up? </p>",
      "rawMarkdown": "rohanrao What would you recommend replacing pandas with to speed it up?",
      "votes": null
    },
    {
      "id": "1066724",
      "postDate": "11/02/2020 00:58:44",
      "content": "<p>Depending on the type of operations, you could use datatable or rapids or bigquery or plain numpy arrays / dictionaries as <a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> mentioned <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193276#1062047\" target=\"_blank\">here</a>.</p>\n<p>Its also worth exploring C (wrapped in Python) for certain processing to get a speedup (though I feel 9hrs is plenty time if code is structured well).</p>\n<p>I'm still working on optimizing my own pipeline, my best submission only has 3 engineered features at the moment.</p>",
      "rawMarkdown": "Depending on the type of operations, you could use datatable or rapids or bigquery or plain numpy arrays / dictionaries as @abhimanyud mentioned [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193276#1062047).\n\nIts also worth exploring C (wrapped in Python) for certain processing to get a speedup (though I feel 9hrs is plenty time if code is structured well).\n\nI'm still working on optimizing my own pipeline, my best submission only has 3 engineered features at the moment.",
      "votes": null
    },
    {
      "id": "1066737",
      "postDate": "11/02/2020 01:26:04",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> Only three engineered features, what kinda of sorcery is this hahahaha</p>",
      "rawMarkdown": "rohanrao Only three engineered features, what kinda of sorcery is this hahahaha",
      "votes": null
    },
    {
      "id": "1066772",
      "postDate": "11/02/2020 02:56:16",
      "content": "<p>Plenty of scores above mine so certainly need to add more 🙂</p>",
      "rawMarkdown": "Plenty of scores above mine so certainly need to add more 🙂",
      "votes": null
    },
    {
      "id": "1066814",
      "postDate": "11/02/2020 04:15:12",
      "content": "<blockquote>\n  <p>my best submission only has 3 engineered features at the moment.</p>\n</blockquote>\n<p>OMG and that gets you to 42nd place ATM. 🙏</p>",
      "rawMarkdown": ">my best submission only has 3 engineered features at the moment.\n\nOMG and that gets you to 42nd place ATM. :pray:",
      "votes": null
    },
    {
      "id": "1067908",
      "postDate": "11/02/2020 20:18:10",
      "content": "<p>Did anyone try to simulate in the notebook a submission, in order to estimate how much time it would take on the test set?</p>",
      "rawMarkdown": "Did anyone try to simulate in the notebook a submission, in order to estimate how much time it would take on the test set?",
      "votes": null
    },
    {
      "id": "1068374",
      "postDate": "11/03/2020 10:39:35",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> do you know of any good tutorials for datatable? I've seen your <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">notebook on reading data in</a> and I'm reading the <a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">user guide/API</a>.<br>\nIt seem very quick, but I keep stumbling into issues.</p>\n<p>For example, how do I decrement the values in a column by 1?<br>\nI started by trying this</p>\n<pre><code>train['part'] = train['part'] - 1 \n</code></pre>\n<p>but got the following error</p>\n<pre><code>TypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n</code></pre>\n<p>This is what I have came up with, but it feels like I'm missing something.</p>\n<pre><code># Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n</code></pre>\n<p>On my local machine this takes 3ms vs 31ms in pandas (<code>train['part'] = train['part'] - 1</code>), so still alot quicker.</p>",
      "rawMarkdown": "rohanrao do you know of any good tutorials for datatable? I've seen your [notebook on reading data in](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid) and I'm reading the [user guide/API](https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html).\nIt seem very quick, but I keep stumbling into issues.\n\nFor example, how do I decrement the values in a column by 1?\nI started by trying this\n```\ntrain['part'] = train['part'] - 1 \n```\nbut got the following error\n```\nTypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n```\nThis is what I have came up with, but it feels like I'm missing something.\n```\n# Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n```\nOn my local machine this takes 3ms vs 31ms in pandas (`train['part'] = train['part'] - 1`), so still alot quicker.",
      "votes": null
    },
    {
      "id": "1068381",
      "postDate": "11/03/2020 10:50:04",
      "content": "<p>The datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.</p>\n<p>Till then you could use this forum or the <a href=\"https://github.com/h2oai/datatable/issues\" target=\"_blank\">Github repo</a>, I'm sure folks will be happy to help, me included.</p>\n<p><code>train['part'] = train[:, dt.f.part - 1]</code></p>\n<p>The <a href=\"https://datatable.readthedocs.io/en/latest/manual/f-expressions.html\" target=\"_blank\">f-expression</a> is used for most of these kind of operations.</p>",
      "rawMarkdown": "The datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.\n\nTill then you could use this forum or the [Github repo](https://github.com/h2oai/datatable/issues), I'm sure folks will be happy to help, me included.\n\n`train['part'] = train[:, dt.f.part - 1]`\n\nThe [f-expression](https://datatable.readthedocs.io/en/latest/manual/f-expressions.html) is used for most of these kind of operations.",
      "votes": null
    },
    {
      "id": "1068383",
      "postDate": "11/03/2020 10:52:55",
      "content": "<p>Thank you! I knew that couldn't be the best way!</p>",
      "rawMarkdown": "Thank you! I knew that couldn't be the best way!",
      "votes": null
    },
    {
      "id": "1068837",
      "postDate": "11/03/2020 19:35:53",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> can I ask you what model you use to get such a huge score without big FE? Btw thanks for the datable trick it speeds up a lot!</p>",
      "rawMarkdown": "rohanrao can I ask you what model you use to get such a huge score without big FE? Btw thanks for the datable trick it speeds up a lot!",
      "votes": null
    },
    {
      "id": "1069312",
      "postDate": "11/04/2020 10:02:14",
      "content": "<p>the loop which make prediction on the hidden set may occupy bigger proportion in running time. In my case,it should cost no more than 5s in the provided 120 samples</p>",
      "rawMarkdown": "the loop which make prediction on the hidden set may occupy bigger proportion in running time. In my case,it should cost no more than 5s in the provided 120 samples",
      "votes": null
    },
    {
      "id": "1069318",
      "postDate": "11/04/2020 10:09:11",
      "content": "<p>agree!hah,it depend more on the time it take to make prediction</p>",
      "rawMarkdown": "agree!hah,it depend more on the time it take to make prediction",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1066311,
      "author_name": "kamalnaithani",
      "author_url": "",
      "post_date": "11/01/2020 15:47:09",
      "content": "<p>It will take time in hours. Data in test data set is huge</p>",
      "votes": null,
      "replies": [
        {
          "id": 1066362,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/01/2020 17:03:35",
          "content": "<p>Not necessarily. It depends on features and structure of code (eg: I don't use pandas, it is super slow)<br>\nMine takes ~ 12 minutes for a successful submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066373,
          "author_name": "kamalnaithani",
          "author_url": "",
          "post_date": "11/01/2020 17:15:39",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> , Thanks for suggestion. As of now i am using pandas only and it is too slow 😏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066391,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/01/2020 17:41:51",
          "content": "<p>I'm around 7hrs and I'm using pandas. Currently trying to move to something faster, I haven't figure out what to change to yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066688,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/01/2020 23:55:53",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> What would you recommend replacing pandas with to speed it up? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066724,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/02/2020 00:58:44",
          "content": "<p>Depending on the type of operations, you could use datatable or rapids or bigquery or plain numpy arrays / dictionaries as <a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> mentioned <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193276#1062047\" target=\"_blank\">here</a>.</p>\n<p>Its also worth exploring C (wrapped in Python) for certain processing to get a speedup (though I feel 9hrs is plenty time if code is structured well).</p>\n<p>I'm still working on optimizing my own pipeline, my best submission only has 3 engineered features at the moment.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066737,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/02/2020 01:26:04",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> Only three engineered features, what kinda of sorcery is this hahahaha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066772,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/02/2020 02:56:16",
          "content": "<p>Plenty of scores above mine so certainly need to add more 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066814,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/02/2020 04:15:12",
          "content": "<blockquote>\n  <p>my best submission only has 3 engineered features at the moment.</p>\n</blockquote>\n<p>OMG and that gets you to 42nd place ATM. 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068374,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/03/2020 10:39:35",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> do you know of any good tutorials for datatable? I've seen your <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">notebook on reading data in</a> and I'm reading the <a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">user guide/API</a>.<br>\nIt seem very quick, but I keep stumbling into issues.</p>\n<p>For example, how do I decrement the values in a column by 1?<br>\nI started by trying this</p>\n<pre><code>train['part'] = train['part'] - 1 \n</code></pre>\n<p>but got the following error</p>\n<pre><code>TypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n</code></pre>\n<p>This is what I have came up with, but it feels like I'm missing something.</p>\n<pre><code># Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n</code></pre>\n<p>On my local machine this takes 3ms vs 31ms in pandas (<code>train['part'] = train['part'] - 1</code>), so still alot quicker.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068381,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/03/2020 10:50:04",
          "content": "<p>The datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.</p>\n<p>Till then you could use this forum or the <a href=\"https://github.com/h2oai/datatable/issues\" target=\"_blank\">Github repo</a>, I'm sure folks will be happy to help, me included.</p>\n<p><code>train['part'] = train[:, dt.f.part - 1]</code></p>\n<p>The <a href=\"https://datatable.readthedocs.io/en/latest/manual/f-expressions.html\" target=\"_blank\">f-expression</a> is used for most of these kind of operations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068383,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/03/2020 10:52:55",
          "content": "<p>Thank you! I knew that couldn't be the best way!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068837,
          "author_name": "matthieuplante",
          "author_url": "",
          "post_date": "11/03/2020 19:35:53",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> can I ask you what model you use to get such a huge score without big FE? Btw thanks for the datable trick it speeds up a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1066683,
      "author_name": "raghuramr",
      "author_url": "",
      "post_date": "11/01/2020 23:49:54",
      "content": "<blockquote>\n  <p>Mine has been running since last 30mins.</p>\n</blockquote>\n<p>Don't worry, I got successful submission after 2 to 3 hours of execution as well depending on models. 😊</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1067908,
      "author_name": "barkay",
      "author_url": "",
      "post_date": "11/02/2020 20:18:10",
      "content": "<p>Did anyone try to simulate in the notebook a submission, in order to estimate how much time it would take on the test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1069318,
          "author_name": "huangtaogan",
          "author_url": "",
          "post_date": "11/04/2020 10:09:11",
          "content": "<p>agree!hah,it depend more on the time it take to make prediction</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1069312,
      "author_name": "huangtaogan",
      "author_url": "",
      "post_date": "11/04/2020 10:02:14",
      "content": "<p>the loop which make prediction on the hidden set may occupy bigger proportion in running time. In my case,it should cost no more than 5s in the provided 120 samples</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1066247": "Mine has been running since last 30mins.",
    "1066311": "It will take time in hours. Data in test data set is huge",
    "1066362": "Not necessarily. It depends on features and structure of code (eg: I don't use pandas, it is super slow)\nMine takes ~ 12 minutes for a successful submission.",
    "1066373": "rohanrao , Thanks for suggestion. As of now i am using pandas only and it is too slow 😏",
    "1066391": "I'm around 7hrs and I'm using pandas. Currently trying to move to something faster, I haven't figure out what to change to yet.",
    "1066683": "> Mine has been running since last 30mins.\n\nDon't worry, I got successful submission after 2 to 3 hours of execution as well depending on models. 😊",
    "1066688": "rohanrao What would you recommend replacing pandas with to speed it up?",
    "1066724": "Depending on the type of operations, you could use datatable or rapids or bigquery or plain numpy arrays / dictionaries as @abhimanyud mentioned [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193276#1062047).\n\nIts also worth exploring C (wrapped in Python) for certain processing to get a speedup (though I feel 9hrs is plenty time if code is structured well).\n\nI'm still working on optimizing my own pipeline, my best submission only has 3 engineered features at the moment.",
    "1066737": "rohanrao Only three engineered features, what kinda of sorcery is this hahahaha",
    "1066772": "Plenty of scores above mine so certainly need to add more 🙂",
    "1066814": ">my best submission only has 3 engineered features at the moment.\n\nOMG and that gets you to 42nd place ATM. :pray:",
    "1067908": "Did anyone try to simulate in the notebook a submission, in order to estimate how much time it would take on the test set?",
    "1068374": "rohanrao do you know of any good tutorials for datatable? I've seen your [notebook on reading data in](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid) and I'm reading the [user guide/API](https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html).\nIt seem very quick, but I keep stumbling into issues.\n\nFor example, how do I decrement the values in a column by 1?\nI started by trying this\n```\ntrain['part'] = train['part'] - 1 \n```\nbut got the following error\n```\nTypeError: unsupported operand type(s) for -: 'datatable.Frame' and 'int'\n```\nThis is what I have came up with, but it feels like I'm missing something.\n```\n# Create a dummy column of 1s\ntrain['ones'] = 1\n\n# Decrement a column\ntrain[:, dt.update(part = dt.f.part - dt.f.ones)]\n\n# Drop the unrequired columns\ntrain = train[:, dt.f[:].remove([dt.f.ones])]\n```\nOn my local machine this takes 3ms vs 31ms in pandas (`train['part'] = train['part'] - 1`), so still alot quicker.",
    "1068381": "The datatable documentation is not great I admit. But there will be more examples and tutorials with the release of datatable 1.0.0 (still some bugs to resolve but can expect it soon). I'm not aware of any really good and comprehensive tutorial yet but I do plan to write one once 1.0.0 is released.\n\nTill then you could use this forum or the [Github repo](https://github.com/h2oai/datatable/issues), I'm sure folks will be happy to help, me included.\n\n`train['part'] = train[:, dt.f.part - 1]`\n\nThe [f-expression](https://datatable.readthedocs.io/en/latest/manual/f-expressions.html) is used for most of these kind of operations.",
    "1068383": "Thank you! I knew that couldn't be the best way!",
    "1068837": "rohanrao can I ask you what model you use to get such a huge score without big FE? Btw thanks for the datable trick it speeds up a lot!",
    "1069312": "the loop which make prediction on the hidden set may occupy bigger proportion in running time. In my case,it should cost no more than 5s in the provided 120 samples",
    "1069318": "agree!hah,it depend more on the time it take to make prediction"
  },
  "source": "meta"
}