{
  "id": 209603,
  "title": "How did you handle software design/engineering?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209603",
  "author_name": "",
  "post_date": "2021-01-08T01:32:58.708147600Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thank you to Kaggle/Riiid for hosting this competition. I think Kaggle did a great job at designing an inference workflow to closely simulate a real-life scenario and prevent any data leakage.</p>\n<p>On the other hand, this meant submitting exclusively through Kaggle Notebooks. I found this to be a big friction point as it was difficult to build well-designed and modular code. I tried to import modules through a Dataset, but management and debugging became too much overhead. In the end, I submitted a pretty rickety wall of 600+ lines of code.</p>\n<p>I'm interested to hear how you all handled this, especially if you are working on teams!</p>",
  "messages": [
    {
      "id": "1143642",
      "postDate": "01/08/2021 01:32:58",
      "content": "<p>Thank you to Kaggle/Riiid for hosting this competition. I think Kaggle did a great job at designing an inference workflow to closely simulate a real-life scenario and prevent any data leakage.</p>\n<p>On the other hand, this meant submitting exclusively through Kaggle Notebooks. I found this to be a big friction point as it was difficult to build well-designed and modular code. I tried to import modules through a Dataset, but management and debugging became too much overhead. In the end, I submitted a pretty rickety wall of 600+ lines of code.</p>\n<p>I'm interested to hear how you all handled this, especially if you are working on teams!</p>",
      "rawMarkdown": "Thank you to Kaggle/Riiid for hosting this competition. I think Kaggle did a great job at designing an inference workflow to closely simulate a real-life scenario and prevent any data leakage.\n\nOn the other hand, this meant submitting exclusively through Kaggle Notebooks. I found this to be a big friction point as it was difficult to build well-designed and modular code. I tried to import modules through a Dataset, but management and debugging became too much overhead. In the end, I submitted a pretty rickety wall of 600+ lines of code.\n\nI'm interested to hear how you all handled this, especially if you are working on teams!",
      "votes": null
    },
    {
      "id": "1143716",
      "postDate": "01/08/2021 02:57:43",
      "content": "<p>I used different codes to create training and inference data. It's very difficult to produce error free codes without testing.</p>\n<p>I created a 40 user data(full data), split the last 10 task containers of 20 users as test data, the remaining as train data.<br>\nThen I split test data into 10 test_df.<br>\nI created features with the full data and train data, and update train data with test_df with inference codes.<br>\nThen I can check every data point in the inference Dataset is same as my full data.</p>\n<p>This testing codes could be written quickly compared to 9 hour waits.</p>",
      "rawMarkdown": "I used different codes to create training and inference data. It's very difficult to produce error free codes without testing.\n\nI created a 40 user data(full data), split the last 10 task containers of 20 users as test data, the remaining as train data.\nThen I split test data into 10 test_df.\nI created features with the full data and train data, and update train data with test_df with inference codes.\nThen I can check every data point in the inference Dataset is same as my full data.\n\nThis testing codes could be written quickly compared to 9 hour waits.",
      "votes": null
    },
    {
      "id": "1143814",
      "postDate": "01/08/2021 04:35:33",
      "content": "<p>If you don't mind, I have a few questions:</p>\n<ul>\n<li>So how many \"rows\" of test data did you have?</li>\n<li>What do you mean by <em>10 test_df</em>?</li>\n<li>What sort of data-structure is test_df?</li>\n<li>Did you use git or any other sort of version control outside of Kaggle Notebooks?</li>\n<li>Which platform did you use for development? For example: AWS, GoogleCloud, home computer, etc…</li>\n</ul>\n<p>Thanks!</p>",
      "rawMarkdown": "If you don't mind, I have a few questions:\n\n* So how many \"rows\" of test data did you have?\n* What do you mean by *10 test_df*?\n* What sort of data-structure is test_df?\n* Did you use git or any other sort of version control outside of Kaggle Notebooks?\n* Which platform did you use for development? For example: AWS, GoogleCloud, home computer, etc...\n\nThanks!",
      "votes": null
    },
    {
      "id": "1143816",
      "postDate": "01/08/2021 04:36:21",
      "content": "<p>For a particular FE, we had .3M pkl files during inference 😅…</p>",
      "rawMarkdown": "For a particular FE, we had .3M pkl files during inference 😅...",
      "votes": null
    },
    {
      "id": "1143836",
      "postDate": "01/08/2021 05:01:03",
      "content": "<p>Each test_df is around 20 rows to simulate iter_test().<br>\nTake a look at this great notebook by tito <br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator</a></p>",
      "rawMarkdown": "Each test_df is around 20 rows to simulate iter_test().\nTake a look at this great notebook by tito \n[https://www.kaggle.com/its7171/time-series-api-iter-test-emulator](url)",
      "votes": null
    },
    {
      "id": "1144948",
      "postDate": "01/08/2021 19:11:06",
      "content": "<p>I agree with this. Debugging generic \"scoring errors\" in total darkness was especially painful. This is how programming with punch cards must have felt back in the days.</p>\n<p>Personally it helped a lot to use the <strong>exact same code</strong> to build user history during pre-preocess and on inference. I even split <code>train.csv</code> in groups and delay each group's <code>user_answer</code> and <code>answered_correctly</code> to the next to simulate the same exact input as in inference.</p>\n<p>Initially I preprocessed data in pandas and inferred using iterative code. This was unmaintainable.</p>",
      "rawMarkdown": "I agree with this. Debugging generic \"scoring errors\" in total darkness was especially painful. This is how programming with punch cards must have felt back in the days.\n\nPersonally it helped a lot to use the **exact same code** to build user history during pre-preocess and on inference. I even split `train.csv` in groups and delay each group's `user_answer` and `answered_correctly` to the next to simulate the same exact input as in inference.\n\nInitially I preprocessed data in pandas and inferred using iterative code. This was unmaintainable.",
      "votes": null
    },
    {
      "id": "1145004",
      "postDate": "01/08/2021 19:56:52",
      "content": "<p>Interesting. I think feeding train in the same manner as iter_test() was a good call.</p>\n<p>I didn't do that, so I had to update certain dictionaries (like the ones containing answered_correctly information) differently in train and infer. I think this was a source of a bug in one of my last minute features that I had to scratch.</p>\n<p>Which platform did you use to develop? I was working on an Amazon EC2 instance and used Github. I found that copying/pasting the code over to Kaggle Notebooks was a particular friction point - it would of been very nice to clone/pull the repository and easily call some sort of main() function.</p>",
      "rawMarkdown": "Interesting. I think feeding train in the same manner as iter_test() was a good call.\n\nI didn't do that, so I had to update certain dictionaries (like the ones containing answered_correctly information) differently in train and infer. I think this was a source of a bug in one of my last minute features that I had to scratch.\n\nWhich platform did you use to develop? I was working on an Amazon EC2 instance and used Github. I found that copying/pasting the code over to Kaggle Notebooks was a particular friction point - it would of been very nice to clone/pull the repository and easily call some sort of main() function.",
      "votes": null
    },
    {
      "id": "1145131",
      "postDate": "01/08/2021 23:19:35",
      "content": "<p>I worked locally as long as I was alone. I had to upgrade my computer, though (RAM and CPU), to keep up with the models as they grew more demanding. </p>\n<p>After my teammate joined, we started using a private github repo to share code. We created new notebooks for each experiment to prevent merges as much as we can. <code>git commit; git push</code> and <code>git stash; git pull --rebase; git stash pop</code> is basically all you need.</p>\n<p>If merging jupyter notebooks is really necessary, <a href=\"https://github.com/jupyter/nbdime\" target=\"_blank\">nbdime</a> is really useful for both merging and diffing notebooks.</p>\n<p>Also checkout the officlal <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">kaggle API tool</a>. This command will upload a kernel, link it to your dataset, enable GPU and run it:</p>\n<pre><code>kaggle k push\n</code></pre>\n<p>You can similarly upload new versions of your dataset, take a look at <code>kaggle datasets -h</code>. This tool is very useful and we had our uploads fully automated.</p>\n<p>I also recommend <a href=\"https://pypi.org/project/ipynb-py-convert/\" target=\"_blank\">ipynb-py-convert</a> to convert from ipynb to py and vice versa. Raw .py scripts run faster and with less RAM overhead than .ipynbs, so our last solutions are all .py scripts.</p>\n<p>Also .ipynb  .py conversion is useful because sometimes it's better to use jupyter (for incremental programming, interactive manipulation, etc). But some other times it's much better to run your code from vscode (to debug long cells, code wrapped in for loops, etc.).</p>\n<p>I also used %lprun to profile the inference loop and reveal bottlenecks. Brief tutorial <a href=\"https://jakevdp.github.io/PythonDataScienceHandbook/01.07-timing-and-profiling.html\" target=\"_blank\">here</a>.</p>\n<p>Phew, I'm realizing we did use a lot of helper tools. I hope this helps anyone :)</p>",
      "rawMarkdown": "I worked locally as long as I was alone. I had to upgrade my computer, though (RAM and CPU), to keep up with the models as they grew more demanding. \n\nAfter my teammate joined, we started using a private github repo to share code. We created new notebooks for each experiment to prevent merges as much as we can. `git commit; git push` and `git stash; git pull --rebase; git stash pop` is basically all you need.\n\nIf merging jupyter notebooks is really necessary, [nbdime](https://github.com/jupyter/nbdime) is really useful for both merging and diffing notebooks.\n\nAlso checkout the officlal [kaggle API tool](https://github.com/Kaggle/kaggle-api). This command will upload a kernel, link it to your dataset, enable GPU and run it:\n\n```\nkaggle k push\n```\n\nYou can similarly upload new versions of your dataset, take a look at `kaggle datasets -h`. This tool is very useful and we had our uploads fully automated.\n\nI also recommend [ipynb-py-convert](https://pypi.org/project/ipynb-py-convert/) to convert from ipynb to py and vice versa. Raw .py scripts run faster and with less RAM overhead than .ipynbs, so our last solutions are all .py scripts.\n\nAlso .ipynb <-> .py conversion is useful because sometimes it's better to use jupyter (for incremental programming, interactive manipulation, etc). But some other times it's much better to run your code from vscode (to debug long cells, code wrapped in for loops, etc.).\n\nI also used %lprun to profile the inference loop and reveal bottlenecks. Brief tutorial [here](https://jakevdp.github.io/PythonDataScienceHandbook/01.07-timing-and-profiling.html).\n\nPhew, I'm realizing we did use a lot of helper tools. I hope this helps anyone :)",
      "votes": null
    },
    {
      "id": "1153492",
      "postDate": "01/14/2021 23:54:49",
      "content": "<p>Apologies for the late reply. But thank you for this, and thank you for the tool recommendations. I will have to check out the Kaggle API in closer detail and %lprun.</p>",
      "rawMarkdown": "Apologies for the late reply. But thank you for this, and thank you for the tool recommendations. I will have to check out the Kaggle API in closer detail and %lprun.",
      "votes": null
    },
    {
      "id": "1155232",
      "postDate": "01/16/2021 10:10:25",
      "content": "<p>Codes pushed to github can be automatically synced to kaggle notebook using github actions.</p>\n<p><a href=\"https://github.com/marketplace/actions/push-kaggle-kernel\" target=\"_blank\">https://github.com/marketplace/actions/push-kaggle-kernel</a></p>\n<p>It saved me a lot of time :)</p>",
      "rawMarkdown": "Codes pushed to github can be automatically synced to kaggle notebook using github actions.\n\nhttps://github.com/marketplace/actions/push-kaggle-kernel\n\n It saved me a lot of time :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143716,
      "author_name": "frankw",
      "author_url": "",
      "post_date": "01/08/2021 02:57:43",
      "content": "<p>I used different codes to create training and inference data. It's very difficult to produce error free codes without testing.</p>\n<p>I created a 40 user data(full data), split the last 10 task containers of 20 users as test data, the remaining as train data.<br>\nThen I split test data into 10 test_df.<br>\nI created features with the full data and train data, and update train data with test_df with inference codes.<br>\nThen I can check every data point in the inference Dataset is same as my full data.</p>\n<p>This testing codes could be written quickly compared to 9 hour waits.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143814,
          "author_name": "npa02012",
          "author_url": "",
          "post_date": "01/08/2021 04:35:33",
          "content": "<p>If you don't mind, I have a few questions:</p>\n<ul>\n<li>So how many \"rows\" of test data did you have?</li>\n<li>What do you mean by <em>10 test_df</em>?</li>\n<li>What sort of data-structure is test_df?</li>\n<li>Did you use git or any other sort of version control outside of Kaggle Notebooks?</li>\n<li>Which platform did you use for development? For example: AWS, GoogleCloud, home computer, etc…</li>\n</ul>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143836,
          "author_name": "frankw",
          "author_url": "",
          "post_date": "01/08/2021 05:01:03",
          "content": "<p>Each test_df is around 20 rows to simulate iter_test().<br>\nTake a look at this great notebook by tito <br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143816,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/08/2021 04:36:21",
      "content": "<p>For a particular FE, we had .3M pkl files during inference 😅…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144948,
      "author_name": "bacterio",
      "author_url": "",
      "post_date": "01/08/2021 19:11:06",
      "content": "<p>I agree with this. Debugging generic \"scoring errors\" in total darkness was especially painful. This is how programming with punch cards must have felt back in the days.</p>\n<p>Personally it helped a lot to use the <strong>exact same code</strong> to build user history during pre-preocess and on inference. I even split <code>train.csv</code> in groups and delay each group's <code>user_answer</code> and <code>answered_correctly</code> to the next to simulate the same exact input as in inference.</p>\n<p>Initially I preprocessed data in pandas and inferred using iterative code. This was unmaintainable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145004,
          "author_name": "npa02012",
          "author_url": "",
          "post_date": "01/08/2021 19:56:52",
          "content": "<p>Interesting. I think feeding train in the same manner as iter_test() was a good call.</p>\n<p>I didn't do that, so I had to update certain dictionaries (like the ones containing answered_correctly information) differently in train and infer. I think this was a source of a bug in one of my last minute features that I had to scratch.</p>\n<p>Which platform did you use to develop? I was working on an Amazon EC2 instance and used Github. I found that copying/pasting the code over to Kaggle Notebooks was a particular friction point - it would of been very nice to clone/pull the repository and easily call some sort of main() function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145131,
          "author_name": "bacterio",
          "author_url": "",
          "post_date": "01/08/2021 23:19:35",
          "content": "<p>I worked locally as long as I was alone. I had to upgrade my computer, though (RAM and CPU), to keep up with the models as they grew more demanding. </p>\n<p>After my teammate joined, we started using a private github repo to share code. We created new notebooks for each experiment to prevent merges as much as we can. <code>git commit; git push</code> and <code>git stash; git pull --rebase; git stash pop</code> is basically all you need.</p>\n<p>If merging jupyter notebooks is really necessary, <a href=\"https://github.com/jupyter/nbdime\" target=\"_blank\">nbdime</a> is really useful for both merging and diffing notebooks.</p>\n<p>Also checkout the officlal <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">kaggle API tool</a>. This command will upload a kernel, link it to your dataset, enable GPU and run it:</p>\n<pre><code>kaggle k push\n</code></pre>\n<p>You can similarly upload new versions of your dataset, take a look at <code>kaggle datasets -h</code>. This tool is very useful and we had our uploads fully automated.</p>\n<p>I also recommend <a href=\"https://pypi.org/project/ipynb-py-convert/\" target=\"_blank\">ipynb-py-convert</a> to convert from ipynb to py and vice versa. Raw .py scripts run faster and with less RAM overhead than .ipynbs, so our last solutions are all .py scripts.</p>\n<p>Also .ipynb  .py conversion is useful because sometimes it's better to use jupyter (for incremental programming, interactive manipulation, etc). But some other times it's much better to run your code from vscode (to debug long cells, code wrapped in for loops, etc.).</p>\n<p>I also used %lprun to profile the inference loop and reveal bottlenecks. Brief tutorial <a href=\"https://jakevdp.github.io/PythonDataScienceHandbook/01.07-timing-and-profiling.html\" target=\"_blank\">here</a>.</p>\n<p>Phew, I'm realizing we did use a lot of helper tools. I hope this helps anyone :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1153492,
          "author_name": "npa02012",
          "author_url": "",
          "post_date": "01/14/2021 23:54:49",
          "content": "<p>Apologies for the late reply. But thank you for this, and thank you for the tool recommendations. I will have to check out the Kaggle API in closer detail and %lprun.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1155232,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/16/2021 10:10:25",
          "content": "<p>Codes pushed to github can be automatically synced to kaggle notebook using github actions.</p>\n<p><a href=\"https://github.com/marketplace/actions/push-kaggle-kernel\" target=\"_blank\">https://github.com/marketplace/actions/push-kaggle-kernel</a></p>\n<p>It saved me a lot of time :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143642": "Thank you to Kaggle/Riiid for hosting this competition. I think Kaggle did a great job at designing an inference workflow to closely simulate a real-life scenario and prevent any data leakage.\n\nOn the other hand, this meant submitting exclusively through Kaggle Notebooks. I found this to be a big friction point as it was difficult to build well-designed and modular code. I tried to import modules through a Dataset, but management and debugging became too much overhead. In the end, I submitted a pretty rickety wall of 600+ lines of code.\n\nI'm interested to hear how you all handled this, especially if you are working on teams!",
    "1143716": "I used different codes to create training and inference data. It's very difficult to produce error free codes without testing.\n\nI created a 40 user data(full data), split the last 10 task containers of 20 users as test data, the remaining as train data.\nThen I split test data into 10 test_df.\nI created features with the full data and train data, and update train data with test_df with inference codes.\nThen I can check every data point in the inference Dataset is same as my full data.\n\nThis testing codes could be written quickly compared to 9 hour waits.",
    "1143814": "If you don't mind, I have a few questions:\n\n* So how many \"rows\" of test data did you have?\n* What do you mean by *10 test_df*?\n* What sort of data-structure is test_df?\n* Did you use git or any other sort of version control outside of Kaggle Notebooks?\n* Which platform did you use for development? For example: AWS, GoogleCloud, home computer, etc...\n\nThanks!",
    "1143816": "For a particular FE, we had .3M pkl files during inference 😅...",
    "1143836": "Each test_df is around 20 rows to simulate iter_test().\nTake a look at this great notebook by tito \n[https://www.kaggle.com/its7171/time-series-api-iter-test-emulator](url)",
    "1144948": "I agree with this. Debugging generic \"scoring errors\" in total darkness was especially painful. This is how programming with punch cards must have felt back in the days.\n\nPersonally it helped a lot to use the **exact same code** to build user history during pre-preocess and on inference. I even split `train.csv` in groups and delay each group's `user_answer` and `answered_correctly` to the next to simulate the same exact input as in inference.\n\nInitially I preprocessed data in pandas and inferred using iterative code. This was unmaintainable.",
    "1145004": "Interesting. I think feeding train in the same manner as iter_test() was a good call.\n\nI didn't do that, so I had to update certain dictionaries (like the ones containing answered_correctly information) differently in train and infer. I think this was a source of a bug in one of my last minute features that I had to scratch.\n\nWhich platform did you use to develop? I was working on an Amazon EC2 instance and used Github. I found that copying/pasting the code over to Kaggle Notebooks was a particular friction point - it would of been very nice to clone/pull the repository and easily call some sort of main() function.",
    "1145131": "I worked locally as long as I was alone. I had to upgrade my computer, though (RAM and CPU), to keep up with the models as they grew more demanding. \n\nAfter my teammate joined, we started using a private github repo to share code. We created new notebooks for each experiment to prevent merges as much as we can. `git commit; git push` and `git stash; git pull --rebase; git stash pop` is basically all you need.\n\nIf merging jupyter notebooks is really necessary, [nbdime](https://github.com/jupyter/nbdime) is really useful for both merging and diffing notebooks.\n\nAlso checkout the officlal [kaggle API tool](https://github.com/Kaggle/kaggle-api). This command will upload a kernel, link it to your dataset, enable GPU and run it:\n\n```\nkaggle k push\n```\n\nYou can similarly upload new versions of your dataset, take a look at `kaggle datasets -h`. This tool is very useful and we had our uploads fully automated.\n\nI also recommend [ipynb-py-convert](https://pypi.org/project/ipynb-py-convert/) to convert from ipynb to py and vice versa. Raw .py scripts run faster and with less RAM overhead than .ipynbs, so our last solutions are all .py scripts.\n\nAlso .ipynb <-> .py conversion is useful because sometimes it's better to use jupyter (for incremental programming, interactive manipulation, etc). But some other times it's much better to run your code from vscode (to debug long cells, code wrapped in for loops, etc.).\n\nI also used %lprun to profile the inference loop and reveal bottlenecks. Brief tutorial [here](https://jakevdp.github.io/PythonDataScienceHandbook/01.07-timing-and-profiling.html).\n\nPhew, I'm realizing we did use a lot of helper tools. I hope this helps anyone :)",
    "1153492": "Apologies for the late reply. But thank you for this, and thank you for the tool recommendations. I will have to check out the Kaggle API in closer detail and %lprun.",
    "1155232": "Codes pushed to github can be automatically synced to kaggle notebook using github actions.\n\nhttps://github.com/marketplace/actions/push-kaggle-kernel\n\n It saved me a lot of time :)"
  },
  "source": "meta"
}