{
  "id": 190354,
  "title": "Incremental / online learning",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190354",
  "author_name": "",
  "post_date": "2020-10-11T11:29:57.578930500Z",
  "votes": 25,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Shameless plug on my notebook trying out some incremental learning with this data<br>\n<a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme</a><br>\nwhich is meant to spark conversation on the upsides and downsides of this approach.</p>\n<p>As far as I'm aware, it's likely not possible to get the same score as with the classical approach where a model is trained offline and deployed statically. However, I feel like using online learning would be very beneficial to this specific problem considering how quickly models like this can be deployed and adapted - it's a breeze to change some features or switch to a different type of ML model as needed. And in the long term, a statically trained model cannot pick up slow changes in the data whereas incremental models are constantly adapting and changing parameters to account for \"data drift\", let's say.</p>\n<p>With good feature engineering, the performance of online learning shouldn't be much worse than that of offline-trained static models and the small accuracy cost might be well worth the reduced effort in maintenance.</p>\n<p>Just wanted to bring this up to see what others think, and draw attention to such different approaches which may be quite useful in real life but are unlikely to ever win a Kaggle competition because they score slightly worse than the typical solutions which are bruteforce trained on a 10 million row timeseries dataset and not necessarily useful in the long run.</p>",
  "messages": [
    {
      "id": "1046137",
      "postDate": "10/11/2020 11:29:57",
      "content": "<p>Shameless plug on my notebook trying out some incremental learning with this data<br>\n<a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme</a><br>\nwhich is meant to spark conversation on the upsides and downsides of this approach.</p>\n<p>As far as I'm aware, it's likely not possible to get the same score as with the classical approach where a model is trained offline and deployed statically. However, I feel like using online learning would be very beneficial to this specific problem considering how quickly models like this can be deployed and adapted - it's a breeze to change some features or switch to a different type of ML model as needed. And in the long term, a statically trained model cannot pick up slow changes in the data whereas incremental models are constantly adapting and changing parameters to account for \"data drift\", let's say.</p>\n<p>With good feature engineering, the performance of online learning shouldn't be much worse than that of offline-trained static models and the small accuracy cost might be well worth the reduced effort in maintenance.</p>\n<p>Just wanted to bring this up to see what others think, and draw attention to such different approaches which may be quite useful in real life but are unlikely to ever win a Kaggle competition because they score slightly worse than the typical solutions which are bruteforce trained on a 10 million row timeseries dataset and not necessarily useful in the long run.</p>",
      "rawMarkdown": "Shameless plug on my notebook trying out some incremental learning with this data\nhttps://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\nwhich is meant to spark conversation on the upsides and downsides of this approach.\n\nAs far as I'm aware, it's likely not possible to get the same score as with the classical approach where a model is trained offline and deployed statically. However, I feel like using online learning would be very beneficial to this specific problem considering how quickly models like this can be deployed and adapted - it's a breeze to change some features or switch to a different type of ML model as needed. And in the long term, a statically trained model cannot pick up slow changes in the data whereas incremental models are constantly adapting and changing parameters to account for \"data drift\", let's say.\n\nWith good feature engineering, the performance of online learning shouldn't be much worse than that of offline-trained static models and the small accuracy cost might be well worth the reduced effort in maintenance.\n\nJust wanted to bring this up to see what others think, and draw attention to such different approaches which may be quite useful in real life but are unlikely to ever win a Kaggle competition because they score slightly worse than the typical solutions which are bruteforce trained on a 10 million row timeseries dataset and not necessarily useful in the long run.",
      "votes": null
    },
    {
      "id": "1046186",
      "postDate": "10/11/2020 12:23:16",
      "content": "<p>for this challenge, it may be a valuable approach.</p>\n<p>e.g. you track a user_id and find that his answer correctness improves over time. then, you may want to give a higher prediction value if you encounter him again in the future</p>",
      "rawMarkdown": "for this challenge, it may be a valuable approach.\n\ne.g. you track a user_id and find that his answer correctness improves over time. then, you may want to give a higher prediction value if you encounter him again in the future",
      "votes": null
    },
    {
      "id": "1046271",
      "postDate": "10/11/2020 14:12:47",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> This can't be done in some sequential ways right?<br>\nCorrect me if I'm wrong..</p>",
      "rawMarkdown": "hengck23 This can't be done in some sequential ways right?\nCorrect me if I'm wrong..",
      "votes": null
    },
    {
      "id": "1046278",
      "postDate": "10/11/2020 14:16:16",
      "content": "<p>You could compare a user's recent answer correctness to their overall answer correctness, this would work sequentially if that's what you mean.</p>",
      "rawMarkdown": "You could compare a user's recent answer correctness to their overall answer correctness, this would work sequentially if that's what you mean.",
      "votes": null
    },
    {
      "id": "1046292",
      "postDate": "10/11/2020 14:27:17",
      "content": "<p>Thanks for bringing this up. Was just wondering do you have more references with you which we can follow up with?<br>\nThanks!</p>",
      "rawMarkdown": "Thanks for bringing this up. Was just wondering do you have more references with you which we can follow up with?\nThanks!",
      "votes": null
    },
    {
      "id": "1046311",
      "postDate": "10/11/2020 14:41:24",
      "content": "<p>Not directly, but the <code>creme</code> documentation (link in my notebook) has some!</p>",
      "rawMarkdown": "Not directly, but the <code>creme</code> documentation (link in my notebook) has some!",
      "votes": null
    },
    {
      "id": "1046320",
      "postDate": "10/11/2020 14:47:19",
      "content": "<p>nvm, this also is a great read incase someone is interested. <a href=\"https://freecontent.manning.com/active-transfer-learning-with-pytorch/\" target=\"_blank\">https://freecontent.manning.com/active-transfer-learning-with-pytorch/</a></p>",
      "rawMarkdown": "nvm, this also is a great read incase someone is interested. https://freecontent.manning.com/active-transfer-learning-with-pytorch/",
      "votes": null
    },
    {
      "id": "1046331",
      "postDate": "10/11/2020 14:53:53",
      "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> the sequential way i meant was like transformer model or lstm model and all</p>",
      "rawMarkdown": "spacelx the sequential way i meant was like transformer model or lstm model and all",
      "votes": null
    },
    {
      "id": "1048007",
      "postDate": "10/13/2020 05:55:11",
      "content": "<p>Just as an update, a bug in the submission loop is now finally fixed and we have a first score of 0.740 - not bad at all, considering that the demonstration kernel only goes through a fraction of the test set and very basic mean correctness features!</p>",
      "rawMarkdown": "Just as an update, a bug in the submission loop is now finally fixed and we have a first score of 0.740 - not bad at all, considering that the demonstration kernel only goes through a fraction of the test set and very basic mean correctness features!",
      "votes": null
    },
    {
      "id": "1048090",
      "postDate": "10/13/2020 07:19:57",
      "content": "<p>Thank you for sharing. Did you find out what the error was in previous versions? … I see that added \".values \" .. only this was an error?</p>",
      "rawMarkdown": "Thank you for sharing. Did you find out what the error was in previous versions? ... I see that added \".values \" .. only this was an error?",
      "votes": null
    },
    {
      "id": "1048109",
      "postDate": "10/13/2020 07:55:17",
      "content": "<p>That was a stupid oversight, the actual error was assuming that <code>prior_group_answers_correct</code> had the length of the previous <code>test_df</code> without lectures, when actually it has the length of <code>test_df</code> in its original state… so it just threw a size mismatch whenever a <code>test_df</code> included lectures.</p>",
      "rawMarkdown": "That was a stupid oversight, the actual error was assuming that `prior_group_answers_correct` had the length of the previous `test_df` without lectures, when actually it has the length of `test_df` in its original state... so it just threw a size mismatch whenever a `test_df` included lectures.",
      "votes": null
    },
    {
      "id": "1048211",
      "postDate": "10/13/2020 09:34:50",
      "content": "<p>Your notebook is a great source of motivation to learn active learning, thanks ! May I ask why you only infer on a fraction of the test set ? Is it related to time or RAM issues ? <br>\nIt looks like more promising than you thought :)</p>",
      "rawMarkdown": "Your notebook is a great source of motivation to learn active learning, thanks ! May I ask why you only infer on a fraction of the test set ? Is it related to time or RAM issues ? \nIt looks like more promising than you thought :)",
      "votes": null
    },
    {
      "id": "1048239",
      "postDate": "10/13/2020 10:12:52",
      "content": "<p>He trained for a fraction only as 100M rows is a lot… But made predictions on all IMO.</p>",
      "rawMarkdown": "He trained for a fraction only as 100M rows is a lot... But made predictions on all IMO.",
      "votes": null
    },
    {
      "id": "1048253",
      "postDate": "10/13/2020 10:31:55",
      "content": "<p>Thanks, I hope it's useful! There are certainly a lot of things to improve and think about with this approach…<br>\nI train on a small section of the train set (~5%) because I'm too impatient to wait for it to go through the whole set - also it might end up hitting the 9 hour limit during inference on the private test set which I want to avoid. The code definitely goes through all the test data though (there are only 3 groups in the public set, that's why you can't see more here).<br>\nI just didn't think it necessary to go through all of the train set since it's just a demonstration.</p>",
      "rawMarkdown": "Thanks, I hope it's useful! There are certainly a lot of things to improve and think about with this approach...\nI train on a small section of the train set (~5%) because I'm too impatient to wait for it to go through the whole set - also it might end up hitting the 9 hour limit during inference on the private test set which I want to avoid. The code definitely goes through all the test data though (there are only 3 groups in the public set, that's why you can't see more here).\nI just didn't think it necessary to go through all of the train set since it's just a demonstration.",
      "votes": null
    },
    {
      "id": "1048272",
      "postDate": "10/13/2020 10:47:12",
      "content": "<p>Also you can increase the batch size and use GPUS and separate the model training from model predicting as well; Lot of work!</p>",
      "rawMarkdown": "Also you can increase the batch size and use GPUS and separate the model training from model predicting as well; Lot of work!",
      "votes": null
    },
    {
      "id": "1048322",
      "postDate": "10/13/2020 11:39:58",
      "content": "<p>Ok it makes sense, I thought you submit on only a portion of the test set</p>",
      "rawMarkdown": "Ok it makes sense, I thought you submit on only a portion of the test set",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1046186,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/11/2020 12:23:16",
      "content": "<p>for this challenge, it may be a valuable approach.</p>\n<p>e.g. you track a user_id and find that his answer correctness improves over time. then, you may want to give a higher prediction value if you encounter him again in the future</p>",
      "votes": null,
      "replies": [
        {
          "id": 1046271,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "10/11/2020 14:12:47",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> This can't be done in some sequential ways right?<br>\nCorrect me if I'm wrong..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046278,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/11/2020 14:16:16",
          "content": "<p>You could compare a user's recent answer correctness to their overall answer correctness, this would work sequentially if that's what you mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046331,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "10/11/2020 14:53:53",
          "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> the sequential way i meant was like transformer model or lstm model and all</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046292,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "10/11/2020 14:27:17",
      "content": "<p>Thanks for bringing this up. Was just wondering do you have more references with you which we can follow up with?<br>\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1046311,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/11/2020 14:41:24",
          "content": "<p>Not directly, but the <code>creme</code> documentation (link in my notebook) has some!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046320,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/11/2020 14:47:19",
          "content": "<p>nvm, this also is a great read incase someone is interested. <a href=\"https://freecontent.manning.com/active-transfer-learning-with-pytorch/\" target=\"_blank\">https://freecontent.manning.com/active-transfer-learning-with-pytorch/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1048007,
      "author_name": "spacelx",
      "author_url": "",
      "post_date": "10/13/2020 05:55:11",
      "content": "<p>Just as an update, a bug in the submission loop is now finally fixed and we have a first score of 0.740 - not bad at all, considering that the demonstration kernel only goes through a fraction of the test set and very basic mean correctness features!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1048090,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "10/13/2020 07:19:57",
          "content": "<p>Thank you for sharing. Did you find out what the error was in previous versions? … I see that added \".values \" .. only this was an error?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048109,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/13/2020 07:55:17",
          "content": "<p>That was a stupid oversight, the actual error was assuming that <code>prior_group_answers_correct</code> had the length of the previous <code>test_df</code> without lectures, when actually it has the length of <code>test_df</code> in its original state… so it just threw a size mismatch whenever a <code>test_df</code> included lectures.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048211,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "10/13/2020 09:34:50",
          "content": "<p>Your notebook is a great source of motivation to learn active learning, thanks ! May I ask why you only infer on a fraction of the test set ? Is it related to time or RAM issues ? <br>\nIt looks like more promising than you thought :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048239,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 10:12:52",
          "content": "<p>He trained for a fraction only as 100M rows is a lot… But made predictions on all IMO.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048253,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/13/2020 10:31:55",
          "content": "<p>Thanks, I hope it's useful! There are certainly a lot of things to improve and think about with this approach…<br>\nI train on a small section of the train set (~5%) because I'm too impatient to wait for it to go through the whole set - also it might end up hitting the 9 hour limit during inference on the private test set which I want to avoid. The code definitely goes through all the test data though (there are only 3 groups in the public set, that's why you can't see more here).<br>\nI just didn't think it necessary to go through all of the train set since it's just a demonstration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048272,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 10:47:12",
          "content": "<p>Also you can increase the batch size and use GPUS and separate the model training from model predicting as well; Lot of work!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048322,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "10/13/2020 11:39:58",
          "content": "<p>Ok it makes sense, I thought you submit on only a portion of the test set</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1046137": "Shameless plug on my notebook trying out some incremental learning with this data\nhttps://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\nwhich is meant to spark conversation on the upsides and downsides of this approach.\n\nAs far as I'm aware, it's likely not possible to get the same score as with the classical approach where a model is trained offline and deployed statically. However, I feel like using online learning would be very beneficial to this specific problem considering how quickly models like this can be deployed and adapted - it's a breeze to change some features or switch to a different type of ML model as needed. And in the long term, a statically trained model cannot pick up slow changes in the data whereas incremental models are constantly adapting and changing parameters to account for \"data drift\", let's say.\n\nWith good feature engineering, the performance of online learning shouldn't be much worse than that of offline-trained static models and the small accuracy cost might be well worth the reduced effort in maintenance.\n\nJust wanted to bring this up to see what others think, and draw attention to such different approaches which may be quite useful in real life but are unlikely to ever win a Kaggle competition because they score slightly worse than the typical solutions which are bruteforce trained on a 10 million row timeseries dataset and not necessarily useful in the long run.",
    "1046186": "for this challenge, it may be a valuable approach.\n\ne.g. you track a user_id and find that his answer correctness improves over time. then, you may want to give a higher prediction value if you encounter him again in the future",
    "1046271": "hengck23 This can't be done in some sequential ways right?\nCorrect me if I'm wrong..",
    "1046278": "You could compare a user's recent answer correctness to their overall answer correctness, this would work sequentially if that's what you mean.",
    "1046292": "Thanks for bringing this up. Was just wondering do you have more references with you which we can follow up with?\nThanks!",
    "1046311": "Not directly, but the <code>creme</code> documentation (link in my notebook) has some!",
    "1046320": "nvm, this also is a great read incase someone is interested. https://freecontent.manning.com/active-transfer-learning-with-pytorch/",
    "1046331": "spacelx the sequential way i meant was like transformer model or lstm model and all",
    "1048007": "Just as an update, a bug in the submission loop is now finally fixed and we have a first score of 0.740 - not bad at all, considering that the demonstration kernel only goes through a fraction of the test set and very basic mean correctness features!",
    "1048090": "Thank you for sharing. Did you find out what the error was in previous versions? ... I see that added \".values \" .. only this was an error?",
    "1048109": "That was a stupid oversight, the actual error was assuming that `prior_group_answers_correct` had the length of the previous `test_df` without lectures, when actually it has the length of `test_df` in its original state... so it just threw a size mismatch whenever a `test_df` included lectures.",
    "1048211": "Your notebook is a great source of motivation to learn active learning, thanks ! May I ask why you only infer on a fraction of the test set ? Is it related to time or RAM issues ? \nIt looks like more promising than you thought :)",
    "1048239": "He trained for a fraction only as 100M rows is a lot... But made predictions on all IMO.",
    "1048253": "Thanks, I hope it's useful! There are certainly a lot of things to improve and think about with this approach...\nI train on a small section of the train set (~5%) because I'm too impatient to wait for it to go through the whole set - also it might end up hitting the 9 hour limit during inference on the private test set which I want to avoid. The code definitely goes through all the test data though (there are only 3 groups in the public set, that's why you can't see more here).\nI just didn't think it necessary to go through all of the train set since it's just a demonstration.",
    "1048272": "Also you can increase the batch size and use GPUS and separate the model training from model predicting as well; Lot of work!",
    "1048322": "Ok it makes sense, I thought you submit on only a portion of the test set"
  },
  "source": "meta"
}