{
  "id": 193276,
  "title": "Any sequence models around there?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/193276",
  "author_name": "Claudio Verdú Ruiz",
  "post_date": "2020-10-26T09:01:30.476000",
  "votes": 27,
  "comment_count": 57,
  "views": 0,
  "content": "<p>Hello guys,</p>\n<p>I'm currently developing a sequence model: everything fine and working but the inference phase seems to be exceeding the time limit.</p>\n<p>The thing is: for every new row in every group in the test set, you have to go for the existing user's rows, concat the new row, compute cummulative features, tail it (or pad it), etc. Those who are working on this will understand. Also add the new confirmed user rows with the previous_correct_answers (this is almost free in time though).</p>\n<p>Is there any obvious trick that I'm missing?</p>\n<p>Thank you mates.</p>\n<p>EDIT:</p>\n<p>After 20 days of twerking the code, I finally got my submission with a strong baseline (0.76) in ~2h of inferencing time. I wanted to share what I did for others to be able to reproduce:</p>\n<ul>\n<li>I was joining questions' features on the fly when training the model, and also in inference time. This is specially slow, so what I did is catching these features in a separated run for the last [windows_size] interactions of every user. I only need those last interactions in inference time to append new interactions from the test set.</li>\n<li>Avoid pandas if possible. Given that for every new row you will be creating a whole 2D input, you will have to compute everything you can on numpy directly. To be able to track every column, you can just create a hashmap with columns and their indices. Tailing is also super slow, you can change that for a slice in numpy to really speed it up.</li>\n</ul>\n<p>I still had to drop some features that I'll be recovering now little by little. The whole thing made me improve from ~1.30s to ~0.25s every 100 rows. </p>",
  "messages": [
    {
      "id": 1060504,
      "postDate": "2020-10-26T09:01:30.477Z",
      "content": "<p>Hello guys,</p>\n<p>I'm currently developing a sequence model: everything fine and working but the inference phase seems to be exceeding the time limit.</p>\n<p>The thing is: for every new row in every group in the test set, you have to go for the existing user's rows, concat the new row, compute cummulative features, tail it (or pad it), etc. Those who are working on this will understand. Also add the new confirmed user rows with the previous_correct_answers (this is almost free in time though).</p>\n<p>Is there any obvious trick that I'm missing?</p>\n<p>Thank you mates.</p>\n<p>EDIT:</p>\n<p>After 20 days of twerking the code, I finally got my submission with a strong baseline (0.76) in ~2h of inferencing time. I wanted to share what I did for others to be able to reproduce:</p>\n<ul>\n<li>I was joining questions' features on the fly when training the model, and also in inference time. This is specially slow, so what I did is catching these features in a separated run for the last [windows_size] interactions of every user. I only need those last interactions in inference time to append new interactions from the test set.</li>\n<li>Avoid pandas if possible. Given that for every new row you will be creating a whole 2D input, you will have to compute everything you can on numpy directly. To be able to track every column, you can just create a hashmap with columns and their indices. Tailing is also super slow, you can change that for a slice in numpy to really speed it up.</li>\n</ul>\n<p>I still had to drop some features that I'll be recovering now little by little. The whole thing made me improve from ~1.30s to ~0.25s every 100 rows. </p>",
      "rawMarkdown": "Hello guys,\n\nI'm currently developing a sequence model: everything fine and working but the inference phase seems to be exceeding the time limit.\n\nThe thing is: for every new row in every group in the test set, you have to go for the existing user's rows, concat the new row, compute cummulative features, tail it (or pad it), etc. Those who are working on this will understand. Also add the new confirmed user rows with the previous_correct_answers (this is almost free in time though).\n\nIs there any obvious trick that I'm missing?\n\nThank you mates.\n\nEDIT:\n\nAfter 20 days of twerking the code, I finally got my submission with a strong baseline (0.76) in ~2h of inferencing time. I wanted to share what I did for others to be able to reproduce:\n\n- I was joining questions' features on the fly when training the model, and also in inference time. This is specially slow, so what I did is catching these features in a separated run for the last [windows_size] interactions of every user. I only need those last interactions in inference time to append new interactions from the test set.\n- Avoid pandas if possible. Given that for every new row you will be creating a whole 2D input, you will have to compute everything you can on numpy directly. To be able to track every column, you can just create a hashmap with columns and their indices. Tailing is also super slow, you can change that for a slice in numpy to really speed it up.\n\nI still had to drop some features that I'll be recovering now little by little. The whole thing made me improve from ~1.30s to ~0.25s every 100 rows. ",
      "votes": 27
    },
    {
      "id": 1062674,
      "postDate": "2020-10-28T04:59:24.463Z",
      "content": "<p>Finally got a competitive score with a sequential model -&gt; 77.1 public lb. I am just glad the pipeline is working smooth. <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> It takes my pipeline about an hour + for inference on the entire test set as well. </p>",
      "rawMarkdown": "Finally got a competitive score with a sequential model -> 77.1 public lb. I am just glad the pipeline is working smooth. @abdurrafae It takes my pipeline about an hour + for inference on the entire test set as well. ",
      "votes": 10,
      "replies": [
        {
          "id": 1062681,
          "postDate": "2020-10-28T05:09:20.757Z",
          "content": "<p>Congrats 🥳 ! Quite efficient pipeline you got there! So it's in typical RNN style or it's Transformers?</p>",
          "rawMarkdown": "Congrats 🥳 ! Quite efficient pipeline you got there! So it's in typical RNN style or it's Transformers?",
          "votes": 1,
          "replies": [
            {
              "id": 1062692,
              "postDate": "2020-10-28T05:24:27.547Z",
              "content": "<p>Thanks Aditya! RNNs were overfitting a lot and I just couldn't figure out the right regularization there. Transformers I believe are better in every way to RNNs(including LSTMs and GRUs)  except when it comes to memory constraints of handling very large( &gt; 500 length) sequences.</p>",
              "rawMarkdown": "Thanks Aditya! RNNs were overfitting a lot and I just couldn't figure out the right regularization there. Transformers I believe are better in every way to RNNs(including LSTMs and GRUs)  except when it comes to memory constraints of handling very large( > 500 length) sequences.",
              "votes": 4
            },
            {
              "id": 1062748,
              "postDate": "2020-10-28T06:41:42.957Z",
              "content": "<p>Great to have a transformer working then!</p>",
              "rawMarkdown": "Great to have a transformer working then!"
            }
          ]
        },
        {
          "id": 1062744,
          "postDate": "2020-10-28T06:35:37.857Z",
          "content": "<p>That's really nice, are you trimming the length of sequence to 100. It's taking quiet a lot longer if I don't limit it.</p>",
          "rawMarkdown": "That's really nice, are you trimming the length of sequence to 100. It's taking quiet a lot longer if I don't limit it."
        },
        {
          "id": 1062771,
          "postDate": "2020-10-28T06:58:45.257Z",
          "content": "<p>Some users have thousands of sessions. Therefore the sequence length will have to be limited to a certain threshold</p>",
          "rawMarkdown": "Some users have thousands of sessions. Therefore the sequence length will have to be limited to a certain threshold"
        },
        {
          "id": 1063399,
          "postDate": "2020-10-28T21:03:32.907Z",
          "content": "<p>Congratulations mate! The most impressive part IMO is the inference time. I'm still struggling to make it work in 8 hours. Would you provide any tricks on this?</p>",
          "rawMarkdown": "Congratulations mate! The most impressive part IMO is the inference time. I'm still struggling to make it work in 8 hours. Would you provide any tricks on this?",
          "votes": 1
        },
        {
          "id": 1063548,
          "postDate": "2020-10-29T04:07:16.320Z",
          "content": "<p>That is probably because I am not using the encoder decoder architecture as described in the SAINT paper. I am using a simpler transformer variant to get a quicker inference time. Will update the inference pipeline time taken for SAINT model as well.</p>\n<p>Tips : Ditch pandas.</p>",
          "rawMarkdown": "That is probably because I am not using the encoder decoder architecture as described in the SAINT paper. I am using a simpler transformer variant to get a quicker inference time. Will update the inference pipeline time taken for SAINT model as well.\n\nTips : Ditch pandas.",
          "votes": 4
        },
        {
          "id": 1063661,
          "postDate": "2020-10-29T07:28:27.873Z",
          "content": "<p>I guess that you are not using the information that we get from test-set API about the correct labels to kinda re-fit; otherwise it shouldn't be less than an hour. Or you are and it's still an hour?</p>",
          "rawMarkdown": "I guess that you are not using the information that we get from test-set API about the correct labels to kinda re-fit; otherwise it shouldn't be less than an hour. Or you are and it's still an hour?",
          "votes": 1
        },
        {
          "id": 1064655,
          "postDate": "2020-10-30T11:33:31.353Z",
          "content": "<p>Refitting during inference is a bad idea imho</p>",
          "rawMarkdown": "Refitting during inference is a bad idea imho",
          "votes": 1
        },
        {
          "id": 1064668,
          "postDate": "2020-10-30T11:47:36.823Z",
          "content": "<blockquote>\n  <p>Refitting during inference is a bad idea imho</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> You say that because the test set represents only 3.15% of the entire dataset?</p>",
          "rawMarkdown": "> Refitting during inference is a bad idea imho\n\n@abhimanyud You say that because the test set represents only 3.15% of the entire dataset?"
        },
        {
          "id": 1064741,
          "postDate": "2020-10-30T13:26:33.090Z",
          "content": "<p>My comment was more specific to Neural networks, since incremental data wont make much difference to model performance while being expensive(which becomes really important given the nature of this competition). Other methods like FTRL which are suited for incremental learning do not fall in the same category. </p>",
          "rawMarkdown": "My comment was more specific to Neural networks, since incremental data wont make much difference to model performance while being expensive(which becomes really important given the nature of this competition). Other methods like FTRL which are suited for incremental learning do not fall in the same category. "
        },
        {
          "id": 1065272,
          "postDate": "2020-10-31T05:33:11.167Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1067841,
          "postDate": "2020-11-02T18:49:59.373Z",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Have you monitored the prediction loss / accuracy during training? I got only about 61 % accuracy for training, haven't tested on the validation dataset. But my loss calculation is based on the whole batch (i.e. the whole sequence for each users - randomly selected with a fixed WINDOW size).</p>",
          "rawMarkdown": "@abhimanyud @claverru Have you monitored the prediction loss / accuracy during training? I got only about 61 % accuracy for training, haven't tested on the validation dataset. But my loss calculation is based on the whole batch (i.e. the whole sequence for each users - randomly selected with a fixed WINDOW size).",
          "votes": 1
        },
        {
          "id": 1068064,
          "postDate": "2020-11-03T03:04:39.030Z",
          "content": "<p>Yes accuracy values are low. Possible reasons -</p>\n<ol>\n<li>Accuracy is being calculated on 0 padded targets as well.</li>\n<li>Since the accuracy is being calculated on the prediction of all time steps, its bound to be lower for the earlier time steps where there is very little information for the model to work with  because of causality.</li>\n</ol>",
          "rawMarkdown": "Yes accuracy values are low. Possible reasons -\n1. Accuracy is being calculated on 0 padded targets as well.\n2. Since the accuracy is being calculated on the prediction of all time steps, its bound to be lower for the earlier time steps where there is very little information for the model to work with  because of causality.",
          "votes": 2
        },
        {
          "id": 1068222,
          "postDate": "2020-11-03T07:30:22.250Z",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> holy fk you are first mate congrats!!</p>",
          "rawMarkdown": "@abhimanyud holy fk you are first mate congrats!!",
          "votes": 2
        },
        {
          "id": 1068283,
          "postDate": "2020-11-03T08:36:50.757Z",
          "content": "<p>I am as shocked as you are 😂</p>",
          "rawMarkdown": "I am as shocked as you are 😂",
          "votes": 2
        },
        {
          "id": 1068314,
          "postDate": "2020-11-03T09:13:36.487Z",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> Could you share your local validation - LB relation? It's been a while since my last submission because even if my intuition tells me I'm doing better, my local validation is dropping. I just want to know if they are even in your case with a solid sequence implementation or there is a huge gap.</p>",
          "rawMarkdown": "@abhimanyud Could you share your local validation - LB relation? It's been a while since my last submission because even if my intuition tells me I'm doing better, my local validation is dropping. I just want to know if they are even in your case with a solid sequence implementation or there is a huge gap.",
          "votes": 1
        },
        {
          "id": 1068334,
          "postDate": "2020-11-03T09:43:28.980Z",
          "content": "<p>Validation and lb more or less increase and decrease in sync. However my local AUC is 86.4 (I know, I know 🤕). Its because I was unable to properly propagate padding masks to the loss.</p>",
          "rawMarkdown": "Validation and lb more or less increase and decrease in sync. However my local AUC is 86.4 (I know, I know 🤕). Its because I was unable to properly propagate padding masks to the loss.",
          "votes": 2
        },
        {
          "id": 1068377,
          "postDate": "2020-11-03T10:45:29.240Z",
          "content": "<p>Doesn't this work for you?</p>\n<pre><code>loss_object = tf.keras.losses.SparseCategoricalCrossentropy(\n    from_logits=True, reduction='none')\n\ndef loss_function(real, pred):\n  mask = tf.math.logical_not(tf.math.equal(real, 0))\n  loss_ = loss_object(real, pred)\n\n  mask = tf.cast(mask, dtype=loss_.dtype)\n  loss_ *= mask\n\n  return tf.reduce_sum(loss_)/tf.reduce_sum(mask)\n\n\ndef accuracy_function(real, pred):\n  accuracies = tf.equal(real, tf.argmax(pred, axis=2))\n\n  mask = tf.math.logical_not(tf.math.equal(real, 0))\n  accuracies = tf.math.logical_and(mask, accuracies)\n\n  accuracies = tf.cast(accuracies, dtype=tf.float32)\n  mask = tf.cast(mask, dtype=tf.float32)\n  return tf.reduce_sum(accuracies)/tf.reduce_sum(mask)\n</code></pre>\n<p>(from <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/transformer</a>)</p>\n<p>If you wanted to use AUC just mask it as in the example above.</p>",
          "rawMarkdown": "Doesn't this work for you?\n\n```python\nloss_object = tf.keras.losses.SparseCategoricalCrossentropy(\n    from_logits=True, reduction='none')\n\ndef loss_function(real, pred):\n  mask = tf.math.logical_not(tf.math.equal(real, 0))\n  loss_ = loss_object(real, pred)\n\n  mask = tf.cast(mask, dtype=loss_.dtype)\n  loss_ *= mask\n\n  return tf.reduce_sum(loss_)/tf.reduce_sum(mask)\n\n\ndef accuracy_function(real, pred):\n  accuracies = tf.equal(real, tf.argmax(pred, axis=2))\n\n  mask = tf.math.logical_not(tf.math.equal(real, 0))\n  accuracies = tf.math.logical_and(mask, accuracies)\n\n  accuracies = tf.cast(accuracies, dtype=tf.float32)\n  mask = tf.cast(mask, dtype=tf.float32)\n  return tf.reduce_sum(accuracies)/tf.reduce_sum(mask)\n```\n(from [https://www.tensorflow.org/tutorials/text/transformer](https://www.tensorflow.org/tutorials/text/transformer))\n\nIf you wanted to use AUC just mask it as in the example above.",
          "votes": 1
        },
        {
          "id": 1068379,
          "postDate": "2020-11-03T10:48:56.463Z",
          "content": "<p>It does work. But only if the masks reach the final layer. In my case, I get DenseToDenseSet Operation not implemented error. So I decided to ditch getting too deep into this error and thought lets just work with unmasked losses. </p>",
          "rawMarkdown": "It does work. But only if the masks reach the final layer. In my case, I get DenseToDenseSet Operation not implemented error. So I decided to ditch getting too deep into this error and thought lets just work with unmasked losses. ",
          "votes": 1
        },
        {
          "id": 1068449,
          "postDate": "2020-11-03T12:33:48.317Z",
          "content": "<p>Well mate I'd say you're doing good :)</p>",
          "rawMarkdown": "Well mate I'd say you're doing good :)",
          "votes": 1
        },
        {
          "id": 1071146,
          "postDate": "2020-11-06T14:54:31.647Z",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> would you mind to share how you use encoder only for training and prediction with transformer model? I think decoder would be too slow for inference, so I want to use only encoder also. But in this case, the problem is that I think I need to mask some input places and only predict on those places. Something like masked language modeling as Bert. Unlike encoder decoder approach, only a small portion of input places will be used for training, otherwise masking too many places will confuse the model. Is your approach similar to this?</p>",
          "rawMarkdown": "@abhimanyud would you mind to share how you use encoder only for training and prediction with transformer model? I think decoder would be too slow for inference, so I want to use only encoder also. But in this case, the problem is that I think I need to mask some input places and only predict on those places. Something like masked language modeling as Bert. Unlike encoder decoder approach, only a small portion of input places will be used for training, otherwise masking too many places will confuse the model. Is your approach similar to this?",
          "votes": 3
        },
        {
          "id": 1071161,
          "postDate": "2020-11-06T15:16:38.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Masked language modeling will not be useful in my opinion, because there is no sense of causality there. That is why BERT (utilizes bidirectional context) is useful for generating good word part embeddings but not for language modeling. Causal modeling (of the kind employed in GPT or Universal transformer) should be the way to go. You want your output at every time step to be dependent on the past but not the future. The encoder only approach uses interaction id as query and response as key value in a self attention layer. That's about it.</p>",
          "rawMarkdown": "@yihdarshieh Masked language modeling will not be useful in my opinion, because there is no sense of causality there. That is why BERT (utilizes bidirectional context) is useful for generating good word part embeddings but not for language modeling. Causal modeling (of the kind employed in GPT or Universal transformer) should be the way to go. You want your output at every time step to be dependent on the past but not the future. The encoder only approach uses interaction id as query and response as key value in a self attention layer. That's about it.",
          "votes": 3
        },
        {
          "id": 1071180,
          "postDate": "2020-11-06T15:31:54.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> Thanks! It helps. Actually, we can apply causal masking in encoder while doing MLM, and this will make the encoder only see the previous position while predict the masked places. (Maybe in this case, the name of encoder is not very proper).<br>\nDo you switch to encoder-decoder already, or still build on encoder-only archeticture?</p>",
          "rawMarkdown": "@abhimanyud Thanks! It helps. Actually, we can apply causal masking in encoder while doing MLM, and this will make the encoder only see the previous position while predict the masked places. (Maybe in this case, the name of encoder is not very proper).\nDo you switch to encoder-decoder already, or still build on encoder-only archeticture?"
        },
        {
          "id": 1072254,
          "postDate": "2020-11-08T01:26:21.850Z",
          "content": "<p>So, if we use only ENCODERs, at the output level, you will have output for every time_stamp in the sequence and with only padding masks if any?</p>",
          "rawMarkdown": "So, if we use only ENCODERs, at the output level, you will have output for every time_stamp in the sequence and with only padding masks if any?"
        }
      ]
    },
    {
      "id": 1075033,
      "postDate": "2020-11-11T10:36:34.163Z",
      "content": "<p>I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">NOTEBOOK</a>.</p>",
      "rawMarkdown": "I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. [NOTEBOOK](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public).",
      "votes": 1
    },
    {
      "id": 1062047,
      "postDate": "2020-10-27T14:28:28.443Z",
      "content": "<p>I am working on it too. The best I have come up with is ditching pandas dataframes altogether and using a hash table of numpy arrays.</p>",
      "rawMarkdown": "I am working on it too. The best I have come up with is ditching pandas dataframes altogether and using a hash table of numpy arrays.",
      "votes": 1,
      "replies": [
        {
          "id": 1062065,
          "postDate": "2020-10-27T14:43:23.917Z",
          "content": "<p>Something similar here, I have my users in a table with user_id as key.</p>",
          "rawMarkdown": "Something similar here, I have my users in a table with user_id as key."
        },
        {
          "id": 1062086,
          "postDate": "2020-10-27T15:02:41.063Z",
          "content": "<p>That seems to be the only working solution that I have found as well. Though it still takes about an hour at inference time just to manage the data pipeline. How are you faring in this regard?</p>",
          "rawMarkdown": "That seems to be the only working solution that I have found as well. Though it still takes about an hour at inference time just to manage the data pipeline. How are you faring in this regard?"
        }
      ]
    },
    {
      "id": 1060573,
      "postDate": "2020-10-26T10:37:30.030Z",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> we have a large training dataset, so:</p>\n<ul>\n<li>compute features etc using only training dataset (offline) and train your sequence model</li>\n<li>inference / test using pre-computed rows (and avoid recomputing for every row)</li>\n</ul>\n<p>Another suggestion is to start with fewer features / lesser number of computations at inference time, then measure the time taken and keep adding code while staying below the kernel time limits.</p>",
      "rawMarkdown": "@claverru we have a large training dataset, so:\n- compute features etc using only training dataset (offline) and train your sequence model\n- inference / test using pre-computed rows (and avoid recomputing for every row)\n\nAnother suggestion is to start with fewer features / lesser number of computations at inference time, then measure the time taken and keep adding code while staying below the kernel time limits.",
      "votes": 1,
      "replies": [
        {
          "id": 1060679,
          "postDate": "2020-10-26T13:11:49.513Z",
          "content": "<p>Hello! thanks for answering. I've trained my model asside, and the weights are loaded on the fly when submitting. I tried to precompute some things to be able to drop user interactions older than my windows size. Though this approach is so prune to fail (hard to addecuate). Also I dont have many features, (less than 30), with a relatively small windows size (lower than 100). </p>",
          "rawMarkdown": "Hello! thanks for answering. I've trained my model asside, and the weights are loaded on the fly when submitting. I tried to precompute some things to be able to drop user interactions older than my windows size. Though this approach is so prune to fail (hard to addecuate). Also I dont have many features, (less than 30), with a relatively small windows size (lower than 100). "
        }
      ]
    },
    {
      "id": 1060558,
      "postDate": "2020-10-26T10:15:14.667Z",
      "content": "<p>I am working on it as well, but , and if it is acceptable in performance (both time and score) share it. Edit: no luck :P</p>\n<p>PS: about time, try to see if different options work better (i.e., join vs merge).</p>",
      "rawMarkdown": "I am working on it as well, but ~~I am dealing with an error that gives me repeated indexes. Hopefully I can dedicate some time to it to fix it~~, and if it is acceptable in performance (both time and score) share it. Edit: no luck :P\n\nPS: about time, try to see if different options work better (i.e., join vs merge).",
      "votes": 1,
      "replies": [
        {
          "id": 1060680,
          "postDate": "2020-10-26T13:12:19.410Z",
          "content": "<p>Try .reset_index(drop=True) before that. </p>",
          "rawMarkdown": "Try .reset_index(drop=True) before that. "
        },
        {
          "id": 1060808,
          "postDate": "2020-10-26T14:51:34.977Z",
          "content": "<p>Maybe if you share a snippet I can try to help. No need to be real code, just something reproducible.</p>",
          "rawMarkdown": "Maybe if you share a snippet I can try to help. No need to be real code, just something reproducible."
        }
      ]
    },
    {
      "id": 1060538,
      "postDate": "2020-10-26T09:39:02.560Z",
      "content": "<p>That's why you don't see seq2seq models yet as the tradeoff for using lgbm let's say vs a DL based model as seq2seq is not so easy. The complexity really increases for DL models and might not be needed as it seems people are using lgbms as of now on top of the lb. Plus it might work for old users, but for new users it's gonna take time to warm-up as you won't have content_ids etc and need to pad them etc etc. As compared to lgbm's, they are pretty good in reducing this complexity but track user history efficiently is also complicated. Plus creating a strong lgbm model means quite strong FE's; Hence there's a tradeoff on both sides…</p>",
      "rawMarkdown": "That's why you don't see seq2seq models yet as the tradeoff for using lgbm let's say vs a DL based model as seq2seq is not so easy. The complexity really increases for DL models and might not be needed as it seems people are using lgbms as of now on top of the lb. Plus it might work for old users, but for new users it's gonna take time to warm-up as you won't have content_ids etc and need to pad them etc etc. As compared to lgbm's, they are pretty good in reducing this complexity but track user history efficiently is also complicated. Plus creating a strong lgbm model means quite strong FE's; Hence there's a tradeoff on both sides...",
      "votes": 1,
      "replies": [
        {
          "id": 1060685,
          "postDate": "2020-10-26T13:17:10.023Z",
          "content": "<p>I'm too committed already xD. I don't really want to enter in a super close feature engineering fight to feed a tree. The only way I see to hit gold is by doing something different, and sequences can give that little advantage. If I can get my current model to work, I think it will surpass the 0.76 mark with room for improvement.</p>",
          "rawMarkdown": "I'm too committed already xD. I don't really want to enter in a super close feature engineering fight to feed a tree. The only way I see to hit gold is by doing something different, and sequences can give that little advantage. If I can get my current model to work, I think it will surpass the 0.76 mark with room for improvement.",
          "votes": 1
        },
        {
          "id": 1060842,
          "postDate": "2020-10-26T15:05:49.013Z",
          "content": "<p>Another thought is to use a sequential model to extract features that you then pass to a non sequential model, like a simple neural net or LGBM. While it is hard to infer with a sequential model in this comp, it is easy to train one. </p>",
          "rawMarkdown": "Another thought is to use a sequential model to extract features that you then pass to a non sequential model, like a simple neural net or LGBM. While it is hard to infer with a sequential model in this comp, it is easy to train one. ",
          "votes": 2
        },
        {
          "id": 1061109,
          "postDate": "2020-10-26T18:46:56.847Z",
          "content": "<p>I appreciate your comment. I've got to train the model, and I'd say it's in a good track. At commiting time I just load the data and weights and start the inference. The inference process though is what is driving me crazy. I need to shorten times to be able to fit in the 9 hours. Also, I can't cache too much to not run OOM.</p>",
          "rawMarkdown": "I appreciate your comment. I've got to train the model, and I'd say it's in a good track. At commiting time I just load the data and weights and start the inference. The inference process though is what is driving me crazy. I need to shorten times to be able to fit in the 9 hours. Also, I can't cache too much to not run OOM.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1065406,
      "postDate": "2020-10-31T08:56:58.907Z",
      "content": "<p>I updated the post with some changes I did to have it finally working.</p>",
      "rawMarkdown": "I updated the post with some changes I did to have it finally working.",
      "votes": 2,
      "replies": [
        {
          "id": 1065412,
          "postDate": "2020-10-31T09:04:46.460Z",
          "content": "<p>Excellent tips! I second all of them. Although you might want to change 'twerking' :P</p>",
          "rawMarkdown": "Excellent tips! I second all of them. Although you might want to change 'twerking' :P",
          "votes": 1
        },
        {
          "id": 1065419,
          "postDate": "2020-10-31T09:11:18.360Z",
          "content": "<p>ha! Twerking is always good!</p>",
          "rawMarkdown": "ha! Twerking is always good!",
          "votes": 1
        },
        {
          "id": 1065444,
          "postDate": "2020-10-31T09:40:31.493Z",
          "content": "<p>Just curious, it's independent of the user_id, right?</p>",
          "rawMarkdown": "Just curious, it's independent of the user_id, right?"
        },
        {
          "id": 1065460,
          "postDate": "2020-10-31T10:06:13.510Z",
          "content": "<p>Yes, never used it, I'm betting for a higher percent of new users on the private test set. </p>",
          "rawMarkdown": "Yes, never used it, I'm betting for a higher percent of new users on the private test set. ",
          "votes": 1
        },
        {
          "id": 1065476,
          "postDate": "2020-10-31T10:24:17.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yea no user_id. The sequence of user interactions  can serve as an identifier for the user in a sequence model. </p>",
          "rawMarkdown": "@adityaecdrid Yea no user_id. The sequence of user interactions  can serve as an identifier for the user in a sequence model. ",
          "votes": 1
        },
        {
          "id": 1065544,
          "postDate": "2020-10-31T12:23:43.117Z",
          "content": "<p>Cool! Thanks for the info.<br>\nWould be great, if you can add the info to the CV/LB discussion for your Val scores Vs LB! Ty!</p>",
          "rawMarkdown": "Cool! Thanks for the info.\nWould be great, if you can add the info to the CV/LB discussion for your Val scores Vs LB! Ty!"
        },
        {
          "id": 1065601,
          "postDate": "2020-10-31T13:50:19.003Z",
          "content": "<p>That's tough actually because I am having some issues with propagating masks to my loss layer. As a result all my CV val aucs are in the 90 + range 👀. But at least changes in local CV match with LB.</p>",
          "rawMarkdown": "That's tough actually because I am having some issues with propagating masks to my loss layer. As a result all my CV val aucs are in the 90 + range 👀. But at least changes in local CV match with LB."
        },
        {
          "id": 1065780,
          "postDate": "2020-10-31T18:19:55.973Z",
          "content": "<p>I added a comment on that discussion. My score increases a lot from V to LB since I'm validating only against unseen users.</p>",
          "rawMarkdown": "I added a comment on that discussion. My score increases a lot from V to LB since I'm validating only against unseen users.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1078873,
      "postDate": "2020-11-15T11:46:20.070Z",
      "content": "<p>how to make batch during test inference ?<br>\nmeans when group has two responses for a user how to track two or more sequence for one iter_test()<br>\nwithout predicting one row at a time ?</p>",
      "rawMarkdown": "how to make batch during test inference ?\nmeans when group has two responses for a user how to track two or more sequence for one iter_test()\nwithout predicting one row at a time ?\n"
    },
    {
      "id": 1061579,
      "postDate": "2020-10-27T05:17:12.553Z",
      "content": "<p>Are you using transformer architecture or RNN architecture? RNNs have a longer inference time overall compared to transformers so may be you can shift to transformers. However the issue of dealing with the large test data still remains.</p>",
      "rawMarkdown": "Are you using transformer architecture or RNN architecture? RNNs have a longer inference time overall compared to transformers so may be you can shift to transformers. However the issue of dealing with the large test data still remains.",
      "replies": [
        {
          "id": 1062064,
          "postDate": "2020-10-27T14:42:50.067Z",
          "content": "<p>I'm using a Transformer based architecture, the model inference itself is pretty fast, the problem is cummulating features, concatenating new rows, etc.</p>",
          "rawMarkdown": "I'm using a Transformer based architecture, the model inference itself is pretty fast, the problem is cummulating features, concatenating new rows, etc."
        },
        {
          "id": 1062078,
          "postDate": "2020-10-27T14:59:49.670Z",
          "content": "<p>Are you using the full size model with 512 units and 4 layers? I have been trying to implement it but it doesn't work when I train the model, it doesn't learn anything. If I use lower dimensions I don't get results which are competitive (hovers around ~0.76 for training). </p>",
          "rawMarkdown": "Are you using the full size model with 512 units and 4 layers? I have been trying to implement it but it doesn't work when I train the model, it doesn't learn anything. If I use lower dimensions I don't get results which are competitive (hovers around ~0.76 for training). "
        },
        {
          "id": 1062094,
          "postDate": "2020-10-27T15:09:10.970Z",
          "content": "<p>Though I m just using the given data as inputs with no FE.</p>",
          "rawMarkdown": "Though I m just using the given data as inputs with no FE."
        },
        {
          "id": 1062115,
          "postDate": "2020-10-27T15:29:49.087Z",
          "content": "<p>Well here's what we can do. Keep a hash table with a deque so that all old users will have the sequence in it with a capped max-len and always keep the Len of the queue accessible as well. So now when you get a new user, you can easily pad etc as needed.</p>\n<p>This is what I am trying to develop as well, so just shared. Also I think we don't have to keep it in mem, use an inverted index on top and wrote to a file via seeking to a line-no etc (Just my current state of the mind, and will test it only a handful of users first, so take it with a grain of salt)</p>\n<p>I might be over complicating things for sure, as top scores which people have are from lgbms, NNs etc I believe as of now.</p>",
          "rawMarkdown": "Well here's what we can do. Keep a hash table with a deque so that all old users will have the sequence in it with a capped max-len and always keep the Len of the queue accessible as well. So now when you get a new user, you can easily pad etc as needed.\n\nThis is what I am trying to develop as well, so just shared. Also I think we don't have to keep it in mem, use an inverted index on top and wrote to a file via seeking to a line-no etc (Just my current state of the mind, and will test it only a handful of users first, so take it with a grain of salt)\n\nI might be over complicating things for sure, as top scores which people have are from lgbms, NNs etc I believe as of now.",
          "votes": 2
        },
        {
          "id": 1062238,
          "postDate": "2020-10-27T17:16:29.273Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  I'm using just two Encoder Layers, with 12 as model dimension and 3 heads (pretty small but data dim is also small).</p>",
          "rawMarkdown": "@abdurrafae  I'm using just two Encoder Layers, with 12 as model dimension and 3 heads (pretty small but data dim is also small)."
        },
        {
          "id": 1062326,
          "postDate": "2020-10-27T18:26:46.237Z",
          "content": "<p>I think that would work. Cause the complete model isn't going anywhere. 😄</p>",
          "rawMarkdown": "I think that would work. Cause the complete model isn't going anywhere. 😄"
        },
        {
          "id": 1062368,
          "postDate": "2020-10-27T19:04:34.633Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> how much time would the memory access takes when keeping data on disk using this approach? I think it might become a bottleneck due to timing constraint.</p>",
          "rawMarkdown": "@adityaecdrid how much time would the memory access takes when keeping data on disk using this approach? I think it might become a bottleneck due to timing constraint."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1062674,
      "author_name": "Abhimanyu Dikshit",
      "author_url": "",
      "post_date": "2020-10-28T04:59:24.463000",
      "content": "<p>Finally got a competitive score with a sequential model -&gt; 77.1 public lb. I am just glad the pipeline is working smooth. <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> It takes my pipeline about an hour + for inference on the entire test set as well. </p>",
      "votes": 10,
      "replies": [
        {
          "id": 1062681,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-28T05:09:20.757000",
          "content": "<p>Congrats 🥳 ! Quite efficient pipeline you got there! So it's in typical RNN style or it's Transformers?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1062692,
              "author_name": "Abhimanyu Dikshit",
              "author_url": "",
              "post_date": "2020-10-28T05:24:27.547000",
              "content": "<p>Thanks Aditya! RNNs were overfitting a lot and I just couldn't figure out the right regularization there. Transformers I believe are better in every way to RNNs(including LSTMs and GRUs)  except when it comes to memory constraints of handling very large( &gt; 500 length) sequences.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 1062748,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-10-28T06:41:42.957000",
              "content": "<p>Great to have a transformer working then!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1062744,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-10-28T06:35:37.857000",
          "content": "<p>That's really nice, are you trimming the length of sequence to 100. It's taking quiet a lot longer if I don't limit it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062771,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-28T06:58:45.257000",
          "content": "<p>Some users have thousands of sessions. Therefore the sequence length will have to be limited to a certain threshold</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1063399,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-28T21:03:32.907000",
          "content": "<p>Congratulations mate! The most impressive part IMO is the inference time. I'm still struggling to make it work in 8 hours. Would you provide any tricks on this?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1063548,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-29T04:07:16.320000",
          "content": "<p>That is probably because I am not using the encoder decoder architecture as described in the SAINT paper. I am using a simpler transformer variant to get a quicker inference time. Will update the inference pipeline time taken for SAINT model as well.</p>\n<p>Tips : Ditch pandas.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1063661,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-29T07:28:27.873000",
          "content": "<p>I guess that you are not using the information that we get from test-set API about the correct labels to kinda re-fit; otherwise it shouldn't be less than an hour. Or you are and it's still an hour?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1064655,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-30T11:33:31.353000",
          "content": "<p>Refitting during inference is a bad idea imho</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1064668,
          "author_name": "Nya 🚀",
          "author_url": "",
          "post_date": "2020-10-30T11:47:36.823000",
          "content": "<blockquote>\n  <p>Refitting during inference is a bad idea imho</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> You say that because the test set represents only 3.15% of the entire dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1064741,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-30T13:26:33.090000",
          "content": "<p>My comment was more specific to Neural networks, since incremental data wont make much difference to model performance while being expensive(which becomes really important given the nature of this competition). Other methods like FTRL which are suited for incremental learning do not fall in the same category. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065272,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-31T05:33:11.167000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1067841,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-02T18:49:59.373000",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Have you monitored the prediction loss / accuracy during training? I got only about 61 % accuracy for training, haven't tested on the validation dataset. But my loss calculation is based on the whole batch (i.e. the whole sequence for each users - randomly selected with a fixed WINDOW size).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1068064,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-11-03T03:04:39.030000",
          "content": "<p>Yes accuracy values are low. Possible reasons -</p>\n<ol>\n<li>Accuracy is being calculated on 0 padded targets as well.</li>\n<li>Since the accuracy is being calculated on the prediction of all time steps, its bound to be lower for the earlier time steps where there is very little information for the model to work with  because of causality.</li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1068222,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-03T07:30:22.250000",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> holy fk you are first mate congrats!!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1068283,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-11-03T08:36:50.757000",
          "content": "<p>I am as shocked as you are 😂</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1068314,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-03T09:13:36.487000",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> Could you share your local validation - LB relation? It's been a while since my last submission because even if my intuition tells me I'm doing better, my local validation is dropping. I just want to know if they are even in your case with a solid sequence implementation or there is a huge gap.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1068334,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-11-03T09:43:28.980000",
          "content": "<p>Validation and lb more or less increase and decrease in sync. However my local AUC is 86.4 (I know, I know 🤕). Its because I was unable to properly propagate padding masks to the loss.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1068377,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-03T10:45:29.240000",
          "content": "<p>Doesn't this work for you?</p>\n<pre><code>loss_object = tf.keras.losses.SparseCategoricalCrossentropy(\n    from_logits=True, reduction='none')\n\ndef loss_function(real, pred):\n  mask = tf.math.logical_not(tf.math.equal(real, 0))\n  loss_ = loss_object(real, pred)\n\n  mask = tf.cast(mask, dtype=loss_.dtype)\n  loss_ *= mask\n\n  return tf.reduce_sum(loss_)/tf.reduce_sum(mask)\n\n\ndef accuracy_function(real, pred):\n  accuracies = tf.equal(real, tf.argmax(pred, axis=2))\n\n  mask = tf.math.logical_not(tf.math.equal(real, 0))\n  accuracies = tf.math.logical_and(mask, accuracies)\n\n  accuracies = tf.cast(accuracies, dtype=tf.float32)\n  mask = tf.cast(mask, dtype=tf.float32)\n  return tf.reduce_sum(accuracies)/tf.reduce_sum(mask)\n</code></pre>\n<p>(from <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/transformer</a>)</p>\n<p>If you wanted to use AUC just mask it as in the example above.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1068379,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-11-03T10:48:56.463000",
          "content": "<p>It does work. But only if the masks reach the final layer. In my case, I get DenseToDenseSet Operation not implemented error. So I decided to ditch getting too deep into this error and thought lets just work with unmasked losses. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1068449,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-03T12:33:48.317000",
          "content": "<p>Well mate I'd say you're doing good :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1071146,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-06T14:54:31.647000",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> would you mind to share how you use encoder only for training and prediction with transformer model? I think decoder would be too slow for inference, so I want to use only encoder also. But in this case, the problem is that I think I need to mask some input places and only predict on those places. Something like masked language modeling as Bert. Unlike encoder decoder approach, only a small portion of input places will be used for training, otherwise masking too many places will confuse the model. Is your approach similar to this?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1071161,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-11-06T15:16:38.097000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Masked language modeling will not be useful in my opinion, because there is no sense of causality there. That is why BERT (utilizes bidirectional context) is useful for generating good word part embeddings but not for language modeling. Causal modeling (of the kind employed in GPT or Universal transformer) should be the way to go. You want your output at every time step to be dependent on the past but not the future. The encoder only approach uses interaction id as query and response as key value in a self attention layer. That's about it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1071180,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-06T15:31:54.137000",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> Thanks! It helps. Actually, we can apply causal masking in encoder while doing MLM, and this will make the encoder only see the previous position while predict the masked places. (Maybe in this case, the name of encoder is not very proper).<br>\nDo you switch to encoder-decoder already, or still build on encoder-only archeticture?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1072254,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-08T01:26:21.850000",
          "content": "<p>So, if we use only ENCODERs, at the output level, you will have output for every time_stamp in the sequence and with only padding masks if any?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1075033,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-11-11T10:36:34.163000",
      "content": "<p>I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">NOTEBOOK</a>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062047,
      "author_name": "Abhimanyu Dikshit",
      "author_url": "",
      "post_date": "2020-10-27T14:28:28.443000",
      "content": "<p>I am working on it too. The best I have come up with is ditching pandas dataframes altogether and using a hash table of numpy arrays.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1062065,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-27T14:43:23.917000",
          "content": "<p>Something similar here, I have my users in a table with user_id as key.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062086,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-10-27T15:02:41.063000",
          "content": "<p>That seems to be the only working solution that I have found as well. Though it still takes about an hour at inference time just to manage the data pipeline. How are you faring in this regard?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1060573,
      "author_name": "Sirish Somanchi",
      "author_url": "",
      "post_date": "2020-10-26T10:37:30.030000",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> we have a large training dataset, so:</p>\n<ul>\n<li>compute features etc using only training dataset (offline) and train your sequence model</li>\n<li>inference / test using pre-computed rows (and avoid recomputing for every row)</li>\n</ul>\n<p>Another suggestion is to start with fewer features / lesser number of computations at inference time, then measure the time taken and keep adding code while staying below the kernel time limits.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1060679,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-26T13:11:49.513000",
          "content": "<p>Hello! thanks for answering. I've trained my model asside, and the weights are loaded on the fly when submitting. I tried to precompute some things to be able to drop user interactions older than my windows size. Though this approach is so prune to fail (hard to addecuate). Also I dont have many features, (less than 30), with a relatively small windows size (lower than 100). </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1060558,
      "author_name": "curious1",
      "author_url": "",
      "post_date": "2020-10-26T10:15:14.667000",
      "content": "<p>I am working on it as well, but , and if it is acceptable in performance (both time and score) share it. Edit: no luck :P</p>\n<p>PS: about time, try to see if different options work better (i.e., join vs merge).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1060680,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-26T13:12:19.410000",
          "content": "<p>Try .reset_index(drop=True) before that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060808,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-26T14:51:34.977000",
          "content": "<p>Maybe if you share a snippet I can try to help. No need to be real code, just something reproducible.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1060538,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-26T09:39:02.560000",
      "content": "<p>That's why you don't see seq2seq models yet as the tradeoff for using lgbm let's say vs a DL based model as seq2seq is not so easy. The complexity really increases for DL models and might not be needed as it seems people are using lgbms as of now on top of the lb. Plus it might work for old users, but for new users it's gonna take time to warm-up as you won't have content_ids etc and need to pad them etc etc. As compared to lgbm's, they are pretty good in reducing this complexity but track user history efficiently is also complicated. Plus creating a strong lgbm model means quite strong FE's; Hence there's a tradeoff on both sides…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1060685,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-26T13:17:10.023000",
          "content": "<p>I'm too committed already xD. I don't really want to enter in a super close feature engineering fight to feed a tree. The only way I see to hit gold is by doing something different, and sequences can give that little advantage. If I can get my current model to work, I think it will surpass the 0.76 mark with room for improvement.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1060842,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2020-10-26T15:05:49.013000",
          "content": "<p>Another thought is to use a sequential model to extract features that you then pass to a non sequential model, like a simple neural net or LGBM. While it is hard to infer with a sequential model in this comp, it is easy to train one. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061109,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-26T18:46:56.847000",
          "content": "<p>I appreciate your comment. I've got to train the model, and I'd say it's in a good track. At commiting time I just load the data and weights and start the inference. The inference process though is what is driving me crazy. I need to shorten times to be able to fit in the 9 hours. Also, I can't cache too much to not run OOM.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1065406,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-10-31T08:56:58.907000",
      "content": "<p>I updated the post with some changes I did to have it finally working.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1065412,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-31T09:04:46.460000",
          "content": "<p>Excellent tips! I second all of them. Although you might want to change 'twerking' :P</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065419,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-31T09:11:18.360000",
          "content": "<p>ha! Twerking is always good!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065444,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-31T09:40:31.493000",
          "content": "<p>Just curious, it's independent of the user_id, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065460,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-31T10:06:13.510000",
          "content": "<p>Yes, never used it, I'm betting for a higher percent of new users on the private test set. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065476,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-31T10:24:17.350000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yea no user_id. The sequence of user interactions  can serve as an identifier for the user in a sequence model. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065544,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-31T12:23:43.117000",
          "content": "<p>Cool! Thanks for the info.<br>\nWould be great, if you can add the info to the CV/LB discussion for your Val scores Vs LB! Ty!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065601,
          "author_name": "Abhimanyu Dikshit",
          "author_url": "",
          "post_date": "2020-10-31T13:50:19.003000",
          "content": "<p>That's tough actually because I am having some issues with propagating masks to my loss layer. As a result all my CV val aucs are in the 90 + range 👀. But at least changes in local CV match with LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065780,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-31T18:19:55.973000",
          "content": "<p>I added a comment on that discussion. My score increases a lot from V to LB since I'm validating only against unseen users.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1078873,
      "author_name": "Dhyey_Patel1234",
      "author_url": "",
      "post_date": "2020-11-15T11:46:20.070000",
      "content": "<p>how to make batch during test inference ?<br>\nmeans when group has two responses for a user how to track two or more sequence for one iter_test()<br>\nwithout predicting one row at a time ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061579,
      "author_name": "AbdurRafae",
      "author_url": "",
      "post_date": "2020-10-27T05:17:12.553000",
      "content": "<p>Are you using transformer architecture or RNN architecture? RNNs have a longer inference time overall compared to transformers so may be you can shift to transformers. However the issue of dealing with the large test data still remains.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1062064,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-27T14:42:50.067000",
          "content": "<p>I'm using a Transformer based architecture, the model inference itself is pretty fast, the problem is cummulating features, concatenating new rows, etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062078,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-10-27T14:59:49.670000",
          "content": "<p>Are you using the full size model with 512 units and 4 layers? I have been trying to implement it but it doesn't work when I train the model, it doesn't learn anything. If I use lower dimensions I don't get results which are competitive (hovers around ~0.76 for training). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062094,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-10-27T15:09:10.970000",
          "content": "<p>Though I m just using the given data as inputs with no FE.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062115,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-27T15:29:49.087000",
          "content": "<p>Well here's what we can do. Keep a hash table with a deque so that all old users will have the sequence in it with a capped max-len and always keep the Len of the queue accessible as well. So now when you get a new user, you can easily pad etc as needed.</p>\n<p>This is what I am trying to develop as well, so just shared. Also I think we don't have to keep it in mem, use an inverted index on top and wrote to a file via seeking to a line-no etc (Just my current state of the mind, and will test it only a handful of users first, so take it with a grain of salt)</p>\n<p>I might be over complicating things for sure, as top scores which people have are from lgbms, NNs etc I believe as of now.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1062238,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-10-27T17:16:29.273000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  I'm using just two Encoder Layers, with 12 as model dimension and 3 heads (pretty small but data dim is also small).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062326,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-10-27T18:26:46.237000",
          "content": "<p>I think that would work. Cause the complete model isn't going anywhere. 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062368,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-10-27T19:04:34.633000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> how much time would the memory access takes when keeping data on disk using this approach? I think it might become a bottleneck due to timing constraint.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1060504": "Hello guys,\n\nI'm currently developing a sequence model: everything fine and working but the inference phase seems to be exceeding the time limit.\n\nThe thing is: for every new row in every group in the test set, you have to go for the existing user's rows, concat the new row, compute cummulative features, tail it (or pad it), etc. Those who are working on this will understand. Also add the new confirmed user rows with the previous_correct_answers (this is almost free in time though).\n\nIs there any obvious trick that I'm missing?\n\nThank you mates.\n\nEDIT:\n\nAfter 20 days of twerking the code, I finally got my submission with a strong baseline (0.76) in ~2h of inferencing time. I wanted to share what I did for others to be able to reproduce:\n\n- I was joining questions' features on the fly when training the model, and also in inference time. This is specially slow, so what I did is catching these features in a separated run for the last [windows_size] interactions of every user. I only need those last interactions in inference time to append new interactions from the test set.\n- Avoid pandas if possible. Given that for every new row you will be creating a whole 2D input, you will have to compute everything you can on numpy directly. To be able to track every column, you can just create a hashmap with columns and their indices. Tailing is also super slow, you can change that for a slice in numpy to really speed it up.\n\nI still had to drop some features that I'll be recovering now little by little. The whole thing made me improve from ~1.30s to ~0.25s every 100 rows. ",
    "1062674": "Finally got a competitive score with a sequential model -> 77.1 public lb. I am just glad the pipeline is working smooth. @abdurrafae It takes my pipeline about an hour + for inference on the entire test set as well. ",
    "1075033": "I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. [NOTEBOOK](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public).",
    "1062047": "I am working on it too. The best I have come up with is ditching pandas dataframes altogether and using a hash table of numpy arrays.",
    "1060573": "@claverru we have a large training dataset, so:\n- compute features etc using only training dataset (offline) and train your sequence model\n- inference / test using pre-computed rows (and avoid recomputing for every row)\n\nAnother suggestion is to start with fewer features / lesser number of computations at inference time, then measure the time taken and keep adding code while staying below the kernel time limits.",
    "1060558": "I am working on it as well, but ~~I am dealing with an error that gives me repeated indexes. Hopefully I can dedicate some time to it to fix it~~, and if it is acceptable in performance (both time and score) share it. Edit: no luck :P\n\nPS: about time, try to see if different options work better (i.e., join vs merge).",
    "1060538": "That's why you don't see seq2seq models yet as the tradeoff for using lgbm let's say vs a DL based model as seq2seq is not so easy. The complexity really increases for DL models and might not be needed as it seems people are using lgbms as of now on top of the lb. Plus it might work for old users, but for new users it's gonna take time to warm-up as you won't have content_ids etc and need to pad them etc etc. As compared to lgbm's, they are pretty good in reducing this complexity but track user history efficiently is also complicated. Plus creating a strong lgbm model means quite strong FE's; Hence there's a tradeoff on both sides...",
    "1065406": "I updated the post with some changes I did to have it finally working.",
    "1078873": "how to make batch during test inference ?\nmeans when group has two responses for a user how to track two or more sequence for one iter_test()\nwithout predicting one row at a time ?\n",
    "1061579": "Are you using transformer architecture or RNN architecture? RNNs have a longer inference time overall compared to transformers so may be you can shift to transformers. However the issue of dealing with the large test data still remains."
  }
}