{
  "id": 191751,
  "title": "FTRL with incremental learning",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191751",
  "author_name": "",
  "post_date": "2020-10-18T12:12:11.253785200Z",
  "votes": 53,
  "comment_count": 14,
  "views": 0,
  "content": "<p>This competition is setup with a process where in addition to the original 100M+ rows of training data, batches of <strong>additional labelled data</strong> become available by iterating through the test data and retrieving the corresponding labels. It is the perfect scenario for implementing <strong>incremental learning models</strong> that are likely to boost performance.</p>\n<p>I've shared an end-to-end basic starter notebook for those interested in trying out <strong>FTRL model</strong>: <a href=\"https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\" target=\"_blank\">https://www.kaggle.com/rohanrao/riiid-ftrl-ftw</a></p>\n<p>📒 Uses native <a href=\"https://www.kaggle.com/rohanrao/python-datatable\" target=\"_blank\">Python datatable</a> for data and model<br>\n🔥 Model training using entire 100M rows within 30 seconds<br>\n➕ Incremental model learning using all test data labels<br>\n⚡ Submission scored within 15 minutes on Kaggle<br>\n🎁 Baseline model without feature engineering or hyper-parameter tuning scores 0.74 on LB</p>\n<p>Read more about FTRL: <a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\" target=\"_blank\">https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf</a></p>\n<p>There is a lot of room for improvement. I won't be surprised if the winning models of this competition are incremental learning ones.</p>\n<p><a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">This Pytorch incremental learning notebook</a> by <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> that uses the <a href=\"https://github.com/creme-ml/creme\" target=\"_blank\">creme</a> library is another good place to start.</p>",
  "messages": [
    {
      "id": "1052912",
      "postDate": "10/18/2020 12:12:11",
      "content": "<p>This competition is setup with a process where in addition to the original 100M+ rows of training data, batches of <strong>additional labelled data</strong> become available by iterating through the test data and retrieving the corresponding labels. It is the perfect scenario for implementing <strong>incremental learning models</strong> that are likely to boost performance.</p>\n<p>I've shared an end-to-end basic starter notebook for those interested in trying out <strong>FTRL model</strong>: <a href=\"https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\" target=\"_blank\">https://www.kaggle.com/rohanrao/riiid-ftrl-ftw</a></p>\n<p>📒 Uses native <a href=\"https://www.kaggle.com/rohanrao/python-datatable\" target=\"_blank\">Python datatable</a> for data and model<br>\n🔥 Model training using entire 100M rows within 30 seconds<br>\n➕ Incremental model learning using all test data labels<br>\n⚡ Submission scored within 15 minutes on Kaggle<br>\n🎁 Baseline model without feature engineering or hyper-parameter tuning scores 0.74 on LB</p>\n<p>Read more about FTRL: <a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\" target=\"_blank\">https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf</a></p>\n<p>There is a lot of room for improvement. I won't be surprised if the winning models of this competition are incremental learning ones.</p>\n<p><a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">This Pytorch incremental learning notebook</a> by <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> that uses the <a href=\"https://github.com/creme-ml/creme\" target=\"_blank\">creme</a> library is another good place to start.</p>",
      "rawMarkdown": "This competition is setup with a process where in addition to the original 100M+ rows of training data, batches of **additional labelled data** become available by iterating through the test data and retrieving the corresponding labels. It is the perfect scenario for implementing **incremental learning models** that are likely to boost performance.\n\nI've shared an end-to-end basic starter notebook for those interested in trying out **FTRL model**: https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\n\n📒 Uses native [Python datatable](https://www.kaggle.com/rohanrao/python-datatable) for data and model\n🔥 Model training using entire 100M rows within 30 seconds\n➕ Incremental model learning using all test data labels\n⚡ Submission scored within 15 minutes on Kaggle\n🎁 Baseline model without feature engineering or hyper-parameter tuning scores 0.74 on LB\n\nRead more about FTRL: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\n\nThere is a lot of room for improvement. I won't be surprised if the winning models of this competition are incremental learning ones.\n\n[This Pytorch incremental learning notebook](https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme) by @spacelx that uses the [creme](https://github.com/creme-ml/creme) library is another good place to start.",
      "votes": null
    },
    {
      "id": "1052920",
      "postDate": "10/18/2020 12:29:28",
      "content": "<p>I was writing the one in <a href=\"https://github.com/VowpalWabbit/vowpal_wabbit/wiki\" target=\"_blank\">vw</a>, quite similar to the one i wrote in <a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/75396\" target=\"_blank\">kaggle Malware Comp</a> :) </p>",
      "rawMarkdown": "I was writing the one in [vw](https://github.com/VowpalWabbit/vowpal_wabbit/wiki), quite similar to the one i wrote in [kaggle Malware Comp](https://www.kaggle.com/c/microsoft-malware-prediction/discussion/75396) :)",
      "votes": null
    },
    {
      "id": "1052951",
      "postDate": "10/18/2020 13:02:33",
      "content": "<p>Thanks for linking my notebook (even though it's clearly not done as smartly as yours)!</p>",
      "rawMarkdown": "Thanks for linking my notebook (even though it's clearly not done as smartly as yours)!",
      "votes": null
    },
    {
      "id": "1053115",
      "postDate": "10/18/2020 15:52:47",
      "content": "<p>how does it handle categorical variables?</p>",
      "rawMarkdown": "how does it handle categorical variables?",
      "votes": null
    },
    {
      "id": "1053121",
      "postDate": "10/18/2020 15:57:49",
      "content": "<p>I hope you will share it 🙂<br>\nThe FTRL baseline should be a good benchmark to beat.</p>",
      "rawMarkdown": "I hope you will share it 🙂\nThe FTRL baseline should be a good benchmark to beat.",
      "votes": null
    },
    {
      "id": "1053131",
      "postDate": "10/18/2020 16:10:07",
      "content": "<p>It only handles categorical variables. Even the numeric inputs are treated as categorical.</p>",
      "rawMarkdown": "It only handles categorical variables. Even the numeric inputs are treated as categorical.",
      "votes": null
    },
    {
      "id": "1053985",
      "postDate": "10/19/2020 14:54:36",
      "content": "<p>Author of <a href=\"https://github.com/creme-ml/creme\" target=\"_blank\">creme</a> (a Python library for incremental learning) chiming in.</p>\n<p>Alas, my current opinion is that a LightGBM model trained on the provided training set will outperform a linear model trained on the training set plus the test set. In fact, I'm pretty sure that LightGBM trained on 3% of the training set is enough to outperform an incremental linear model. Of course, I would love to be proven wrong :)</p>",
      "rawMarkdown": "Author of [creme](https://github.com/creme-ml/creme) (a Python library for incremental learning) chiming in.\n\nAlas, my current opinion is that a LightGBM model trained on the provided training set will outperform a linear model trained on the training set plus the test set. In fact, I'm pretty sure that LightGBM trained on 3% of the training set is enough to outperform an incremental linear model. Of course, I would love to be proven wrong :)",
      "votes": null
    },
    {
      "id": "1053996",
      "postDate": "10/19/2020 15:09:23",
      "content": "<p>That's what I would expect as well yeah - especially since I don't expect there to be much drift in this competition's data, neither concerning the average correctness of a user nor the average correctness a question is answered with (undeniably the two strongest features).</p>\n<p>However, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. And I had good fun with that (using the creme package)! Thank you for investing time and effort into making this a nice and accessible package!</p>\n<p>One thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s initial post): I started out by using <code>fit_many</code> to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over <code>fit_one</code> for each row separately. Is there a logical reason for this which I'm not seeing?</p>",
      "rawMarkdown": "That's what I would expect as well yeah - especially since I don't expect there to be much drift in this competition's data, neither concerning the average correctness of a user nor the average correctness a question is answered with (undeniably the two strongest features).\n\nHowever, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. And I had good fun with that (using the creme package)! Thank you for investing time and effort into making this a nice and accessible package!\n\nOne thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of @rohanrao's initial post): I started out by using `fit_many` to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over `fit_one` for each row separately. Is there a logical reason for this which I'm not seeing?",
      "votes": null
    },
    {
      "id": "1054032",
      "postDate": "10/19/2020 15:47:28",
      "content": "<p>Thanks for the kind words Alex!</p>\n<blockquote>\n  <p>However, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. </p>\n</blockquote>\n<p>Exactly! Hopefully with time I'll manage to convince more and more people that this is the case :)</p>\n<blockquote>\n  <p>One thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s initial post): I started out by using fit_many to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over fit_one for each row separately. Is there a logical reason for this which I'm not seeing?</p>\n</blockquote>\n<p>What's happening with <code>fit_many</code> is that the individual gradients are averaged, and the average gradient is used to update the weights. Therefore, using <code>fit_many</code> is not equivalent to <code>fit_one</code>. This might change in the future! Sorry for the confusion.</p>",
      "rawMarkdown": "Thanks for the kind words Alex!\n\n> However, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. \n\nExactly! Hopefully with time I'll manage to convince more and more people that this is the case :)\n\n> One thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of @rohanrao's initial post): I started out by using fit_many to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over fit_one for each row separately. Is there a logical reason for this which I'm not seeing?\n\nWhat's happening with `fit_many` is that the individual gradients are averaged, and the average gradient is used to update the weights. Therefore, using `fit_many` is not equivalent to `fit_one`. This might change in the future! Sorry for the confusion.",
      "votes": null
    },
    {
      "id": "1054041",
      "postDate": "10/19/2020 15:57:47",
      "content": "<p>Ah, thanks for the explanation!</p>",
      "rawMarkdown": "Ah, thanks for the explanation!",
      "votes": null
    },
    {
      "id": "1054185",
      "postDate": "10/19/2020 18:53:40",
      "content": "<p>A side note, Deploying ML model is easy, but Deploying it reliably is hard. Here we have access to test set labels but in real world, it's not that way 😂. We have different sort of metrics etc to determine whether we should consider retrain etc etc! That's why active learning is still kinda new but I feel it hasn't yet received the attention it should be receiving!</p>",
      "rawMarkdown": "A side note, Deploying ML model is easy, but Deploying it reliably is hard. Here we have access to test set labels but in real world, it's not that way 😂. We have different sort of metrics etc to determine whether we should consider retrain etc etc! That's why active learning is still kinda new but I feel it hasn't yet received the attention it should be receiving!",
      "votes": null
    },
    {
      "id": "1054193",
      "postDate": "10/19/2020 18:59:45",
      "content": "<blockquote>\n  <p>Here we have access to test set labels but in real world, it's not that way</p>\n</blockquote>\n<p>We don't have access to test set labels because we are predicting before getting labels. That's how it is in real world too: You predict today and get the labels tomorrow. You can use tomorrow's labels to predict the day after. And so on. Per day or per batch.</p>\n<p>That's why I really like this competition setup. Mimics real world and eases deployment complexity.</p>",
      "rawMarkdown": "> Here we have access to test set labels but in real world, it's not that way\n\nWe don't have access to test set labels because we are predicting before getting labels. That's how it is in real world too: You predict today and get the labels tomorrow. You can use tomorrow's labels to predict the day after. And so on. Per day or per batch.\n\nThat's why I really like this competition setup. Mimics real world and eases deployment complexity.",
      "votes": null
    },
    {
      "id": "1054194",
      "postDate": "10/19/2020 19:02:11",
      "content": "<p>I completely agree with <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s enthusiasm: it's great to see Kaggle competitions move towards a format that is closer to a real-world setting.</p>",
      "rawMarkdown": "I completely agree with @rohanrao's enthusiasm: it's great to see Kaggle competitions move towards a format that is closer to a real-world setting.",
      "votes": null
    },
    {
      "id": "1054475",
      "postDate": "10/20/2020 01:25:59",
      "content": "<p>Yep! It's a nice set-up. My previous comment was kinda poorly worded. Apologies!</p>",
      "rawMarkdown": "Yep! It's a nice set-up. My previous comment was kinda poorly worded. Apologies!",
      "votes": null
    },
    {
      "id": "1623320",
      "postDate": "12/19/2021 18:05:38",
      "content": "<p>I am a novice in this field and I am trying to solve a prediction problem with problems about concept drift. Your example helped me a lot. Do you think this algorithm(ftrl) has the advantage of being used to solve problems that are more time-sensitive? For example, in a dataset containing one year's data, to predict a target for December, only the data closer to December is valid for the model. If this problem is solved in a normal training way, the early data will have side effects. What do you think about this. Thank you.😊</p>",
      "rawMarkdown": "I am a novice in this field and I am trying to solve a prediction problem with problems about concept drift. Your example helped me a lot. Do you think this algorithm(ftrl) has the advantage of being used to solve problems that are more time-sensitive? For example, in a dataset containing one year's data, to predict a target for December, only the data closer to December is valid for the model. If this problem is solved in a normal training way, the early data will have side effects. What do you think about this. Thank you.😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1052920,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "10/18/2020 12:29:28",
      "content": "<p>I was writing the one in <a href=\"https://github.com/VowpalWabbit/vowpal_wabbit/wiki\" target=\"_blank\">vw</a>, quite similar to the one i wrote in <a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/75396\" target=\"_blank\">kaggle Malware Comp</a> :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1053121,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/18/2020 15:57:49",
          "content": "<p>I hope you will share it 🙂<br>\nThe FTRL baseline should be a good benchmark to beat.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1052951,
      "author_name": "spacelx",
      "author_url": "",
      "post_date": "10/18/2020 13:02:33",
      "content": "<p>Thanks for linking my notebook (even though it's clearly not done as smartly as yours)!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1053115,
      "author_name": "alaeddineayadi",
      "author_url": "",
      "post_date": "10/18/2020 15:52:47",
      "content": "<p>how does it handle categorical variables?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1053131,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/18/2020 16:10:07",
          "content": "<p>It only handles categorical variables. Even the numeric inputs are treated as categorical.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1053985,
      "author_name": "maxhalford",
      "author_url": "",
      "post_date": "10/19/2020 14:54:36",
      "content": "<p>Author of <a href=\"https://github.com/creme-ml/creme\" target=\"_blank\">creme</a> (a Python library for incremental learning) chiming in.</p>\n<p>Alas, my current opinion is that a LightGBM model trained on the provided training set will outperform a linear model trained on the training set plus the test set. In fact, I'm pretty sure that LightGBM trained on 3% of the training set is enough to outperform an incremental linear model. Of course, I would love to be proven wrong :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1053996,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/19/2020 15:09:23",
          "content": "<p>That's what I would expect as well yeah - especially since I don't expect there to be much drift in this competition's data, neither concerning the average correctness of a user nor the average correctness a question is answered with (undeniably the two strongest features).</p>\n<p>However, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. And I had good fun with that (using the creme package)! Thank you for investing time and effort into making this a nice and accessible package!</p>\n<p>One thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s initial post): I started out by using <code>fit_many</code> to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over <code>fit_one</code> for each row separately. Is there a logical reason for this which I'm not seeing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054032,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "10/19/2020 15:47:28",
          "content": "<p>Thanks for the kind words Alex!</p>\n<blockquote>\n  <p>However, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. </p>\n</blockquote>\n<p>Exactly! Hopefully with time I'll manage to convince more and more people that this is the case :)</p>\n<blockquote>\n  <p>One thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s initial post): I started out by using fit_many to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over fit_one for each row separately. Is there a logical reason for this which I'm not seeing?</p>\n</blockquote>\n<p>What's happening with <code>fit_many</code> is that the individual gradients are averaged, and the average gradient is used to update the weights. Therefore, using <code>fit_many</code> is not equivalent to <code>fit_one</code>. This might change in the future! Sorry for the confusion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054041,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/19/2020 15:57:47",
          "content": "<p>Ah, thanks for the explanation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054185,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/19/2020 18:53:40",
          "content": "<p>A side note, Deploying ML model is easy, but Deploying it reliably is hard. Here we have access to test set labels but in real world, it's not that way 😂. We have different sort of metrics etc to determine whether we should consider retrain etc etc! That's why active learning is still kinda new but I feel it hasn't yet received the attention it should be receiving!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054193,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/19/2020 18:59:45",
          "content": "<blockquote>\n  <p>Here we have access to test set labels but in real world, it's not that way</p>\n</blockquote>\n<p>We don't have access to test set labels because we are predicting before getting labels. That's how it is in real world too: You predict today and get the labels tomorrow. You can use tomorrow's labels to predict the day after. And so on. Per day or per batch.</p>\n<p>That's why I really like this competition setup. Mimics real world and eases deployment complexity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054194,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "10/19/2020 19:02:11",
          "content": "<p>I completely agree with <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s enthusiasm: it's great to see Kaggle competitions move towards a format that is closer to a real-world setting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054475,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/20/2020 01:25:59",
          "content": "<p>Yep! It's a nice set-up. My previous comment was kinda poorly worded. Apologies!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1623320,
      "author_name": "nevcao",
      "author_url": "",
      "post_date": "12/19/2021 18:05:38",
      "content": "<p>I am a novice in this field and I am trying to solve a prediction problem with problems about concept drift. Your example helped me a lot. Do you think this algorithm(ftrl) has the advantage of being used to solve problems that are more time-sensitive? For example, in a dataset containing one year's data, to predict a target for December, only the data closer to December is valid for the model. If this problem is solved in a normal training way, the early data will have side effects. What do you think about this. Thank you.😊</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1052912": "This competition is setup with a process where in addition to the original 100M+ rows of training data, batches of **additional labelled data** become available by iterating through the test data and retrieving the corresponding labels. It is the perfect scenario for implementing **incremental learning models** that are likely to boost performance.\n\nI've shared an end-to-end basic starter notebook for those interested in trying out **FTRL model**: https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\n\n📒 Uses native [Python datatable](https://www.kaggle.com/rohanrao/python-datatable) for data and model\n🔥 Model training using entire 100M rows within 30 seconds\n➕ Incremental model learning using all test data labels\n⚡ Submission scored within 15 minutes on Kaggle\n🎁 Baseline model without feature engineering or hyper-parameter tuning scores 0.74 on LB\n\nRead more about FTRL: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\n\nThere is a lot of room for improvement. I won't be surprised if the winning models of this competition are incremental learning ones.\n\n[This Pytorch incremental learning notebook](https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme) by @spacelx that uses the [creme](https://github.com/creme-ml/creme) library is another good place to start.",
    "1052920": "I was writing the one in [vw](https://github.com/VowpalWabbit/vowpal_wabbit/wiki), quite similar to the one i wrote in [kaggle Malware Comp](https://www.kaggle.com/c/microsoft-malware-prediction/discussion/75396) :)",
    "1052951": "Thanks for linking my notebook (even though it's clearly not done as smartly as yours)!",
    "1053115": "how does it handle categorical variables?",
    "1053121": "I hope you will share it 🙂\nThe FTRL baseline should be a good benchmark to beat.",
    "1053131": "It only handles categorical variables. Even the numeric inputs are treated as categorical.",
    "1053985": "Author of [creme](https://github.com/creme-ml/creme) (a Python library for incremental learning) chiming in.\n\nAlas, my current opinion is that a LightGBM model trained on the provided training set will outperform a linear model trained on the training set plus the test set. In fact, I'm pretty sure that LightGBM trained on 3% of the training set is enough to outperform an incremental linear model. Of course, I would love to be proven wrong :)",
    "1053996": "That's what I would expect as well yeah - especially since I don't expect there to be much drift in this competition's data, neither concerning the average correctness of a user nor the average correctness a question is answered with (undeniably the two strongest features).\n\nHowever, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. And I had good fun with that (using the creme package)! Thank you for investing time and effort into making this a nice and accessible package!\n\nOne thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of @rohanrao's initial post): I started out by using `fit_many` to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over `fit_one` for each row separately. Is there a logical reason for this which I'm not seeing?",
    "1054032": "Thanks for the kind words Alex!\n\n> However, I see the strength of incremental learning rather in the ease of deployment and upkeep, hence why I wanted to play around with it. \n\nExactly! Hopefully with time I'll manage to convince more and more people that this is the case :)\n\n> One thing I stumbled across - I'm just doing linear regression in my example kernel (see link at the bottom of @rohanrao's initial post): I started out by using fit_many to fit each batch, but found there to be a significant score improvement when replacing this call by a loop over fit_one for each row separately. Is there a logical reason for this which I'm not seeing?\n\nWhat's happening with `fit_many` is that the individual gradients are averaged, and the average gradient is used to update the weights. Therefore, using `fit_many` is not equivalent to `fit_one`. This might change in the future! Sorry for the confusion.",
    "1054041": "Ah, thanks for the explanation!",
    "1054185": "A side note, Deploying ML model is easy, but Deploying it reliably is hard. Here we have access to test set labels but in real world, it's not that way 😂. We have different sort of metrics etc to determine whether we should consider retrain etc etc! That's why active learning is still kinda new but I feel it hasn't yet received the attention it should be receiving!",
    "1054193": "> Here we have access to test set labels but in real world, it's not that way\n\nWe don't have access to test set labels because we are predicting before getting labels. That's how it is in real world too: You predict today and get the labels tomorrow. You can use tomorrow's labels to predict the day after. And so on. Per day or per batch.\n\nThat's why I really like this competition setup. Mimics real world and eases deployment complexity.",
    "1054194": "I completely agree with @rohanrao's enthusiasm: it's great to see Kaggle competitions move towards a format that is closer to a real-world setting.",
    "1054475": "Yep! It's a nice set-up. My previous comment was kinda poorly worded. Apologies!",
    "1623320": "I am a novice in this field and I am trying to solve a prediction problem with problems about concept drift. Your example helped me a lot. Do you think this algorithm(ftrl) has the advantage of being used to solve problems that are more time-sensitive? For example, in a dataset containing one year's data, to predict a target for December, only the data closer to December is valid for the model. If this problem is solved in a normal training way, the early data will have side effects. What do you think about this. Thank you.😊"
  },
  "source": "meta"
}