{
  "id": 205250,
  "title": "To lecture or not to lecture",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205250",
  "author_name": "",
  "post_date": "2020-12-19T07:14:48.141345200Z",
  "votes": 14,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I've been seeing quite varied opinions on usage of lectures in the models. So it'll be nice to know how everyone is faring with these and how much are people gaining (or losing) after addition of lectures.</p>\n<p>My stats are:<br>\nModel : Transfomer-esque (Val: 0.79248, LV: 0.792)<br>\nLectures : Included<br>\nGain/ after addition: ~0.003</p>\n<p>Not sure if the gain was que to randomness or purely due to the addition of lectures as I made some more changes to the architecture and can't compare against the original baseline again. </p>",
  "messages": [
    {
      "id": "1118551",
      "postDate": "12/19/2020 07:14:48",
      "content": "<p>I've been seeing quite varied opinions on usage of lectures in the models. So it'll be nice to know how everyone is faring with these and how much are people gaining (or losing) after addition of lectures.</p>\n<p>My stats are:<br>\nModel : Transfomer-esque (Val: 0.79248, LV: 0.792)<br>\nLectures : Included<br>\nGain/ after addition: ~0.003</p>\n<p>Not sure if the gain was que to randomness or purely due to the addition of lectures as I made some more changes to the architecture and can't compare against the original baseline again. </p>",
      "rawMarkdown": "I've been seeing quite varied opinions on usage of lectures in the models. So it'll be nice to know how everyone is faring with these and how much are people gaining (or losing) after addition of lectures.\n\nMy stats are:\nModel : Transfomer-esque (Val: 0.79248, LV: 0.792)\nLectures : Included\nGain/~~Loss~~ after addition: ~0.003\n\nNot sure if the gain was que to randomness or purely due to the addition of lectures as I made some more changes to the architecture and can't compare against the original baseline again.",
      "votes": null
    },
    {
      "id": "1118597",
      "postDate": "12/19/2020 08:30:28",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> I've tried to use lectures in transformer too like with <code>[Q1, Q2, Q3, L1, Q4, Q5, L2, Q5 ...]</code> with new encoding for <code>Ln</code> to not conflict with <code>Qn</code> but it did not provide any gain. How do you use lectures in sequence? Dedicated embedding?</p>",
      "rawMarkdown": "abdurrafae I've tried to use lectures in transformer too like with `[Q1, Q2, Q3, L1, Q4, Q5, L2, Q5 ...]` with new encoding for `Ln` to not conflict with `Qn` but it did not provide any gain. How do you use lectures in sequence? Dedicated embedding?",
      "votes": null
    },
    {
      "id": "1118612",
      "postDate": "12/19/2020 08:41:25",
      "content": "<p>I encoded both of them using the same layer after ensuring that all content_ids are different for questions and lectures.</p>\n<p>This increases a bit of preprocessing in the pipeline at inference but simplifies the model.</p>",
      "rawMarkdown": "I encoded both of them using the same layer after ensuring that all content_ids are different for questions and lectures.\n\nThis increases a bit of preprocessing in the pipeline at inference but simplifies the model.",
      "votes": null
    },
    {
      "id": "1118631",
      "postDate": "12/19/2020 08:56:04",
      "content": "<p>Thanks for the answer. That's exactly what I've done, and what answer value did you encode for lecture? I made the prediction to be softmax with 3 classes but results were better without lectures.</p>",
      "rawMarkdown": "Thanks for the answer. That's exactly what I've done, and what answer value did you encode for lecture? I made the prediction to be softmax with 3 classes but results were better without lectures.",
      "votes": null
    },
    {
      "id": "1118632",
      "postDate": "12/19/2020 08:58:33",
      "content": "<p>I'm masking the prediction for lectures from my loss function. So model only learns to predict the answer for questions.</p>",
      "rawMarkdown": "I'm masking the prediction for lectures from my loss function. So model only learns to predict the answer for questions.",
      "votes": null
    },
    {
      "id": "1118635",
      "postDate": "12/19/2020 09:01:54",
      "content": "<p>What's your model size and current score for transformer? (I assume your team's score is from LGBM)</p>\n<p>I have 128 d_model and 4 num_layers. Not sure how much help would scaling up the d_model parameter would achieve and I don't have the luxury (GPU resources) for testing it out at the moment. Keeping this change for the last week so I can iterate the current sized model a bit faster. </p>",
      "rawMarkdown": "What's your model size and current score for transformer? (I assume your team's score is from LGBM)\n\nI have 128 d_model and 4 num_layers. Not sure how much help would scaling up the d_model parameter would achieve and I don't have the luxury (GPU resources) for testing it out at the moment. Keeping this change for the last week so I can iterate the current sized model a bit faster.",
      "votes": null
    },
    {
      "id": "1118645",
      "postDate": "12/19/2020 09:15:10",
      "content": "<p>d_model=256 and num_layers=4. We don't have any issue with GPU resources, limitation is more with host RAM. You get OOM on GPU? or you don't have enough GPU quota to train with more data?</p>",
      "rawMarkdown": "d_model=256 and num_layers=4. We don't have any issue with GPU resources, limitation is more with host RAM. You get OOM on GPU? or you don't have enough GPU quota to train with more data?",
      "votes": null
    },
    {
      "id": "1118650",
      "postDate": "12/19/2020 09:17:43",
      "content": "<p>It takes around 20-24 hours to train one model once using Kaggle GPUs. Then I don't have much quota left for training a new one in the same week.</p>",
      "rawMarkdown": "It takes around 20-24 hours to train one model once using Kaggle GPUs. Then I don't have much quota left for training a new one in the same week.",
      "votes": null
    },
    {
      "id": "1118653",
      "postDate": "12/19/2020 09:20:37",
      "content": "<p>Maybe you could try optimize training (dataloader), how long is one epoch? It's 12 minutes for me.</p>",
      "rawMarkdown": "Maybe you could try optimize training (dataloader), how long is one epoch? It's 12 minutes for me.",
      "votes": null
    },
    {
      "id": "1118702",
      "postDate": "12/19/2020 10:12:19",
      "content": "<p>My batch generator (tf.keras) is able to iterate through the whole dataset in less than a minute. So I know the issue is in the long training time of the model. 1 epoch is around 30 mins and I train upto 42 epochs for now.</p>",
      "rawMarkdown": "My batch generator (tf.keras) is able to iterate through the whole dataset in less than a minute. So I know the issue is in the long training time of the model. 1 epoch is around 30 mins and I train upto 42 epochs for now.",
      "votes": null
    },
    {
      "id": "1118703",
      "postDate": "12/19/2020 10:16:41",
      "content": "<p>I see, so your modified architecture should increase a lot the total number of model parameters, like in the way you combinate the embeddings?</p>",
      "rawMarkdown": "I see, so your modified architecture should increase a lot the total number of model parameters, like in the way you combinate the embeddings?",
      "votes": null
    },
    {
      "id": "1118725",
      "postDate": "12/19/2020 10:47:21",
      "content": "<p>The model size isn't that large ~4 - 5 M parameters overall. I think it may be due to my own implementation of attention module and masks. </p>",
      "rawMarkdown": "The model size isn't that large ~4 - 5 M parameters overall. I think it may be due to my own implementation of attention module and masks.",
      "votes": null
    },
    {
      "id": "1119092",
      "postDate": "12/19/2020 17:48:46",
      "content": "<p>Anyone knows why there are some questions and lectures with the same ID? And why, on the other hand, they can have different parts?</p>",
      "rawMarkdown": "Anyone knows why there are some questions and lectures with the same ID? And why, on the other hand, they can have different parts?",
      "votes": null
    },
    {
      "id": "1119097",
      "postDate": "12/19/2020 17:56:29",
      "content": "<p>I think that's just an artifact of the data generating process and the number coincide by chance. I don't think there is any significance of a matching lecture_id and question_id.</p>",
      "rawMarkdown": "I think that's just an artifact of the data generating process and the number coincide by chance. I don't think there is any significance of a matching lecture_id and question_id.",
      "votes": null
    },
    {
      "id": "1119171",
      "postDate": "12/19/2020 19:20:25",
      "content": "<p>May be, yeah… And the parts are the same/equivalent? I'm thinking about integrating lectures on my Riid-based model, but with 2% of appearance is it even worth it?</p>",
      "rawMarkdown": "May be, yeah... And the parts are the same/equivalent? I'm thinking about integrating lectures on my Riid-based model, but with 2% of appearance is it even worth it?",
      "votes": null
    },
    {
      "id": "1119193",
      "postDate": "12/19/2020 19:54:32",
      "content": "<p>I use the lecture from the beginning, so not able to say which is better. I make sure it is not used for loss calculation, and not included in the inference during submitting.</p>",
      "rawMarkdown": "I use the lecture from the beginning, so not able to say which is better. I make sure it is not used for loss calculation, and not included in the inference during submitting.",
      "votes": null
    },
    {
      "id": "1119196",
      "postDate": "12/19/2020 19:59:15",
      "content": "<p>I don't know which framework you use, <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> . I use Google Colab Pro + Tensorflow + TPU -&gt; each epoch (~1.024 M sequences of window size 128) take 6m. Maybe it is late for you to setup TPU, but you should consider this (with Google Colab , pro or not) as an option.</p>",
      "rawMarkdown": "I don't know which framework you use, @abdurrafae . I use Google Colab Pro + Tensorflow + TPU -> each epoch (~1.024 M sequences of window size 128) take 6m. Maybe it is late for you to setup TPU, but you should consider this (with Google Colab , pro or not) as an option.",
      "votes": null
    },
    {
      "id": "1119228",
      "postDate": "12/19/2020 20:37:11",
      "content": "<p>I use TF keras. My dataset can be loaded on to RAM all at once so I think setting up the TPU shouldn't be that hard. I'll definitely try that out for my next model training.</p>\n<p>I also have almost same number of sequences btw :D</p>\n<p>If you don't mind telling, how big is your model as in d_model and num_layers for the 6 min training time? and can you try it out on GPU for one epoch so I can benchmark off of it. </p>",
      "rawMarkdown": "I use TF keras. My dataset can be loaded on to RAM all at once so I think setting up the TPU shouldn't be that hard. I'll definitely try that out for my next model training.\n\nI also have almost same number of sequences btw :D\n\nIf you don't mind telling, how big is your model as in d_model and num_layers for the 6 min training time? and can you try it out on GPU for one epoch so I can benchmark off of it.",
      "votes": null
    },
    {
      "id": "1119239",
      "postDate": "12/19/2020 20:51:10",
      "content": "<p>d_dim = 256, ffn_dim = 1024, n_layers = 4. Don't remember exactly the timing for GPU, but for each batch (128 sequences of window size 128), looks like 0.25 second, while with TPU, the same number of sequences is processed by TPU in 0.04 - 0.05 seconds.</p>\n<p>I use TFRecord files and there are quite a lot transformation - so it might not make the full power of TPUs.</p>",
      "rawMarkdown": "d_dim = 256, ffn_dim = 1024, n_layers = 4. Don't remember exactly the timing for GPU, but for each batch (128 sequences of window size 128), looks like 0.25 second, while with TPU, the same number of sequences is processed by TPU in 0.04 - 0.05 seconds.\n\nI use TFRecord files and there are quite a lot transformation - so it might not make the full power of TPUs.",
      "votes": null
    },
    {
      "id": "1119245",
      "postDate": "12/19/2020 20:54:10",
      "content": "<p>Thanks. Even if it only take me to 15-20 mins, I still save a few hours in each iteration of the model then.</p>",
      "rawMarkdown": "Thanks. Even if it only take me to 15-20 mins, I still save a few hours in each iteration of the model then.",
      "votes": null
    },
    {
      "id": "1119252",
      "postDate": "12/19/2020 21:00:48",
      "content": "<p>That's the entire reason I opened this thread. To get to know it's worth :D</p>",
      "rawMarkdown": "That's the entire reason I opened this thread. To get to know it's worth :D",
      "votes": null
    },
    {
      "id": "1119629",
      "postDate": "12/20/2020 08:43:19",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> I made another try with lectures but it did not provide any gain on CV. I might be doing something wrong. I made sure to encode lectures_id to be after questions_id and mask them in loss. The probability to get lectures in a random sequence is quite low, did you force the selection to have some lectures?</p>",
      "rawMarkdown": "abdurrafae I made another try with lectures but it did not provide any gain on CV. I might be doing something wrong. I made sure to encode lectures_id to be after questions_id and mask them in loss. The probability to get lectures in a random sequence is quite low, did you force the selection to have some lectures?",
      "votes": null
    },
    {
      "id": "1119678",
      "postDate": "12/20/2020 09:30:21",
      "content": "<p>I create my training dataset by taking windows from the original sequences so I'm sure that lectures would be included. </p>\n<p>The thing with Lectures is that the hassle to include them in the pipeline in inference is too much and since I faced a gain when I included them, I'm quite reluctant to roll back. However, I've not sure how much an actual gain they give and wanted to get a feel for that in this thread. </p>\n<p>It's easier to see the gain/importance in LGBM. I wouldn't ask what your features related to lectures are (if any) in your LGBM model, but it'll be really helpful to know if they are important enough or not.</p>",
      "rawMarkdown": "I create my training dataset by taking windows from the original sequences so I'm sure that lectures would be included. \n\nThe thing with Lectures is that the hassle to include them in the pipeline in inference is too much and since I faced a gain when I included them, I'm quite reluctant to roll back. However, I've not sure how much an actual gain they give and wanted to get a feel for that in this thread. \n\nIt's easier to see the gain/importance in LGBM. I wouldn't ask what your features related to lectures are (if any) in your LGBM model, but it'll be really helpful to know if they are important enough or not.",
      "votes": null
    },
    {
      "id": "1119687",
      "postDate": "12/20/2020 09:44:36",
      "content": "<p>Easy to answer, all attempts to use lectures so far in LGBM failed.</p>",
      "rawMarkdown": "Easy to answer, all attempts to use lectures so far in LGBM failed.",
      "votes": null
    },
    {
      "id": "1119688",
      "postDate": "12/20/2020 09:46:38",
      "content": "<p>that is sad</p>",
      "rawMarkdown": "that is sad",
      "votes": null
    },
    {
      "id": "1119692",
      "postDate": "12/20/2020 09:56:49",
      "content": "<p>I woke up this morning with a new model trained with lectures. It did not improve.</p>",
      "rawMarkdown": "I woke up this morning with a new model trained with lectures. It did not improve.",
      "votes": null
    },
    {
      "id": "1119934",
      "postDate": "12/20/2020 13:13:39",
      "content": "<p>Although we shouldn't be emotional, but I still feel a bit disappointed that lectures doesn't help - It means that some very appealing ideas and intuitions don't apply to models - and we need to try a lot of options to find out.</p>\n<p>Actually, at once, I had a feeling to rework on my dataset pipeline and inference pipeline to remove lectures (because I didn't know if including it hurt my score) - but it requires quite a lot of my work, so I didn't change it.</p>",
      "rawMarkdown": "Although we shouldn't be emotional, but I still feel a bit disappointed that lectures doesn't help - It means that some very appealing ideas and intuitions don't apply to models - and we need to try a lot of options to find out.\n\nActually, at once, I had a feeling to rework on my dataset pipeline and inference pipeline to remove lectures (because I didn't know if including it hurt my score) - but it requires quite a lot of my work, so I didn't change it.",
      "votes": null
    },
    {
      "id": "1120249",
      "postDate": "12/20/2020 17:34:39",
      "content": "<p>Will it not work even if we discard lectures that the users didnt really spend enough time? I suppose thru some crude ways we can get the time lecture watched? But the data is so sparse that maybe it is not helping. Intuition-wise it definitely should. I saw one top10er saying he uses lectures..though maybe he has considered it by default in his original design itself<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107993\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107993</a></p>",
      "rawMarkdown": "Will it not work even if we discard lectures that the users didnt really spend enough time? I suppose thru some crude ways we can get the time lecture watched? But the data is so sparse that maybe it is not helping. Intuition-wise it definitely should. I saw one top10er saying he uses lectures..though maybe he has considered it by default in his original design itself\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107993",
      "votes": null
    },
    {
      "id": "1121385",
      "postDate": "12/21/2020 15:59:00",
      "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Can you confirm if we need all datatypes to be converted to int32/64 or float32 for using the TPUs?</p>",
      "rawMarkdown": "yihdarshieh Can you confirm if we need all datatypes to be converted to int32/64 or float32 for using the TPUs?",
      "votes": null
    },
    {
      "id": "1127125",
      "postDate": "12/26/2020 09:10:38",
      "content": "<p>Update:<br>\nI set up a pipeline for data without lectures on the same model that I had fitted data with lectures.</p>\n<p>Model fitted on <strong>data containing lectures</strong> had a better train/val auc score on each epoch. I stopped the test after only 10 epochs though to preserve my quota. (difference in val_auc was around ~.002 by the 10th epoch)</p>",
      "rawMarkdown": "Update:\nI set up a pipeline for data without lectures on the same model that I had fitted data with lectures.\n\nModel fitted on **data containing lectures** had a better train/val auc score on each epoch. I stopped the test after only 10 epochs though to preserve my quota. (difference in val_auc was around ~.002 by the 10th epoch)",
      "votes": null
    },
    {
      "id": "1127133",
      "postDate": "12/26/2020 09:25:08",
      "content": "<p>Great to know lectures help, otherwise the kids won't never study … :)</p>",
      "rawMarkdown": "Great to know lectures help, otherwise the kids won't never study ... :)",
      "votes": null
    },
    {
      "id": "1127137",
      "postDate": "12/26/2020 09:28:59",
      "content": "<p>Excellent, and they would use this thread as proof lectures are not worth. 😆</p>",
      "rawMarkdown": "Excellent, and they would use this thread as proof lectures are not worth. 😆",
      "votes": null
    },
    {
      "id": "1127140",
      "postDate": "12/26/2020 09:30:11",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Sorry for asking in this thread with another question - I saw once, in the early stage of this competition, you tried to used masked LM loss to pretrain the encoder, and said there was a 0.005 gain.</p>\n<p>Now, with your improved model, does masked LM loss pretraining encoder still help?</p>\n<p>I tried it, it gives no gain (even a bit worse, but not much). And when I checked the accuracy of token (i.e <code>content_id</code> of questions / lectures) prediction, the best I can get is about 28% - 34%.</p>\n<p>I really wonder if MLM helps here. </p>",
      "rawMarkdown": "abdurrafae Sorry for asking in this thread with another question - I saw once, in the early stage of this competition, you tried to used masked LM loss to pretrain the encoder, and said there was a 0.005 gain.\n\nNow, with your improved model, does masked LM loss pretraining encoder still help?\n\nI tried it, it gives no gain (even a bit worse, but not much). And when I checked the accuracy of token (i.e `content_id` of questions / lectures) prediction, the best I can get is about 28% - 34%.\n\nI really wonder if MLM helps here.",
      "votes": null
    },
    {
      "id": "1127148",
      "postDate": "12/26/2020 09:44:18",
      "content": "<p>I've only used binary-cross entropy loss with masks for lectures so far. I also remember someone talking about MLM but I didn't understand much about it</p>",
      "rawMarkdown": "I've only used binary-cross entropy loss with masks for lectures so far. I also remember someone talking about MLM but I didn't understand much about it",
      "votes": null
    },
    {
      "id": "1127259",
      "postDate": "12/26/2020 11:27:47",
      "content": "<p>I use the history of lectures. However, I think that if you don't treat them carefully when using them with questions, you will damage the model.</p>",
      "rawMarkdown": "I use the history of lectures. However, I think that if you don't treat them carefully when using them with questions, you will damage the model.",
      "votes": null
    },
    {
      "id": "1130905",
      "postDate": "12/29/2020 12:03:24",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> did you test if the improvement on validation translates to LB?</p>",
      "rawMarkdown": "abdurrafae did you test if the improvement on validation translates to LB?",
      "votes": null
    },
    {
      "id": "1130962",
      "postDate": "12/29/2020 13:02:52",
      "content": "<p>I'm afraid that to do so I'll have to change a lot of my inference pipeline. Hence I didn't check that</p>",
      "rawMarkdown": "I'm afraid that to do so I'll have to change a lot of my inference pipeline. Hence I didn't check that",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1118597,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "12/19/2020 08:30:28",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> I've tried to use lectures in transformer too like with <code>[Q1, Q2, Q3, L1, Q4, Q5, L2, Q5 ...]</code> with new encoding for <code>Ln</code> to not conflict with <code>Qn</code> but it did not provide any gain. How do you use lectures in sequence? Dedicated embedding?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1118612,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 08:41:25",
          "content": "<p>I encoded both of them using the same layer after ensuring that all content_ids are different for questions and lectures.</p>\n<p>This increases a bit of preprocessing in the pipeline at inference but simplifies the model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118631,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/19/2020 08:56:04",
          "content": "<p>Thanks for the answer. That's exactly what I've done, and what answer value did you encode for lecture? I made the prediction to be softmax with 3 classes but results were better without lectures.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118632,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 08:58:33",
          "content": "<p>I'm masking the prediction for lectures from my loss function. So model only learns to predict the answer for questions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118635,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 09:01:54",
          "content": "<p>What's your model size and current score for transformer? (I assume your team's score is from LGBM)</p>\n<p>I have 128 d_model and 4 num_layers. Not sure how much help would scaling up the d_model parameter would achieve and I don't have the luxury (GPU resources) for testing it out at the moment. Keeping this change for the last week so I can iterate the current sized model a bit faster. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118645,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/19/2020 09:15:10",
          "content": "<p>d_model=256 and num_layers=4. We don't have any issue with GPU resources, limitation is more with host RAM. You get OOM on GPU? or you don't have enough GPU quota to train with more data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118650,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 09:17:43",
          "content": "<p>It takes around 20-24 hours to train one model once using Kaggle GPUs. Then I don't have much quota left for training a new one in the same week.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118653,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/19/2020 09:20:37",
          "content": "<p>Maybe you could try optimize training (dataloader), how long is one epoch? It's 12 minutes for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118702,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 10:12:19",
          "content": "<p>My batch generator (tf.keras) is able to iterate through the whole dataset in less than a minute. So I know the issue is in the long training time of the model. 1 epoch is around 30 mins and I train upto 42 epochs for now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118703,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/19/2020 10:16:41",
          "content": "<p>I see, so your modified architecture should increase a lot the total number of model parameters, like in the way you combinate the embeddings?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118725,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 10:47:21",
          "content": "<p>The model size isn't that large ~4 - 5 M parameters overall. I think it may be due to my own implementation of attention module and masks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119196,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/19/2020 19:59:15",
          "content": "<p>I don't know which framework you use, <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> . I use Google Colab Pro + Tensorflow + TPU -&gt; each epoch (~1.024 M sequences of window size 128) take 6m. Maybe it is late for you to setup TPU, but you should consider this (with Google Colab , pro or not) as an option.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119228,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 20:37:11",
          "content": "<p>I use TF keras. My dataset can be loaded on to RAM all at once so I think setting up the TPU shouldn't be that hard. I'll definitely try that out for my next model training.</p>\n<p>I also have almost same number of sequences btw :D</p>\n<p>If you don't mind telling, how big is your model as in d_model and num_layers for the 6 min training time? and can you try it out on GPU for one epoch so I can benchmark off of it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119239,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/19/2020 20:51:10",
          "content": "<p>d_dim = 256, ffn_dim = 1024, n_layers = 4. Don't remember exactly the timing for GPU, but for each batch (128 sequences of window size 128), looks like 0.25 second, while with TPU, the same number of sequences is processed by TPU in 0.04 - 0.05 seconds.</p>\n<p>I use TFRecord files and there are quite a lot transformation - so it might not make the full power of TPUs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119245,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 20:54:10",
          "content": "<p>Thanks. Even if it only take me to 15-20 mins, I still save a few hours in each iteration of the model then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119629,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/20/2020 08:43:19",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> I made another try with lectures but it did not provide any gain on CV. I might be doing something wrong. I made sure to encode lectures_id to be after questions_id and mask them in loss. The probability to get lectures in a random sequence is quite low, did you force the selection to have some lectures?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119678,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/20/2020 09:30:21",
          "content": "<p>I create my training dataset by taking windows from the original sequences so I'm sure that lectures would be included. </p>\n<p>The thing with Lectures is that the hassle to include them in the pipeline in inference is too much and since I faced a gain when I included them, I'm quite reluctant to roll back. However, I've not sure how much an actual gain they give and wanted to get a feel for that in this thread. </p>\n<p>It's easier to see the gain/importance in LGBM. I wouldn't ask what your features related to lectures are (if any) in your LGBM model, but it'll be really helpful to know if they are important enough or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119687,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/20/2020 09:44:36",
          "content": "<p>Easy to answer, all attempts to use lectures so far in LGBM failed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119688,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/20/2020 09:46:38",
          "content": "<p>that is sad</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121385,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/21/2020 15:59:00",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Can you confirm if we need all datatypes to be converted to int32/64 or float32 for using the TPUs?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1119092,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "12/19/2020 17:48:46",
      "content": "<p>Anyone knows why there are some questions and lectures with the same ID? And why, on the other hand, they can have different parts?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119097,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 17:56:29",
          "content": "<p>I think that's just an artifact of the data generating process and the number coincide by chance. I don't think there is any significance of a matching lecture_id and question_id.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119171,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "12/19/2020 19:20:25",
          "content": "<p>May be, yeah… And the parts are the same/equivalent? I'm thinking about integrating lectures on my Riid-based model, but with 2% of appearance is it even worth it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119252,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/19/2020 21:00:48",
          "content": "<p>That's the entire reason I opened this thread. To get to know it's worth :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1119193,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "12/19/2020 19:54:32",
      "content": "<p>I use the lecture from the beginning, so not able to say which is better. I make sure it is not used for loss calculation, and not included in the inference during submitting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1119692,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "12/20/2020 09:56:49",
      "content": "<p>I woke up this morning with a new model trained with lectures. It did not improve.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119934,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/20/2020 13:13:39",
          "content": "<p>Although we shouldn't be emotional, but I still feel a bit disappointed that lectures doesn't help - It means that some very appealing ideas and intuitions don't apply to models - and we need to try a lot of options to find out.</p>\n<p>Actually, at once, I had a feeling to rework on my dataset pipeline and inference pipeline to remove lectures (because I didn't know if including it hurt my score) - but it requires quite a lot of my work, so I didn't change it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120249,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/20/2020 17:34:39",
          "content": "<p>Will it not work even if we discard lectures that the users didnt really spend enough time? I suppose thru some crude ways we can get the time lecture watched? But the data is so sparse that maybe it is not helping. Intuition-wise it definitely should. I saw one top10er saying he uses lectures..though maybe he has considered it by default in his original design itself<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107993\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107993</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127125,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "12/26/2020 09:10:38",
      "content": "<p>Update:<br>\nI set up a pipeline for data without lectures on the same model that I had fitted data with lectures.</p>\n<p>Model fitted on <strong>data containing lectures</strong> had a better train/val auc score on each epoch. I stopped the test after only 10 epochs though to preserve my quota. (difference in val_auc was around ~.002 by the 10th epoch)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1127133,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/26/2020 09:25:08",
          "content": "<p>Great to know lectures help, otherwise the kids won't never study … :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127137,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/26/2020 09:28:59",
          "content": "<p>Excellent, and they would use this thread as proof lectures are not worth. 😆</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127140,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/26/2020 09:30:11",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Sorry for asking in this thread with another question - I saw once, in the early stage of this competition, you tried to used masked LM loss to pretrain the encoder, and said there was a 0.005 gain.</p>\n<p>Now, with your improved model, does masked LM loss pretraining encoder still help?</p>\n<p>I tried it, it gives no gain (even a bit worse, but not much). And when I checked the accuracy of token (i.e <code>content_id</code> of questions / lectures) prediction, the best I can get is about 28% - 34%.</p>\n<p>I really wonder if MLM helps here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127148,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/26/2020 09:44:18",
          "content": "<p>I've only used binary-cross entropy loss with masks for lectures so far. I also remember someone talking about MLM but I didn't understand much about it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130905,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "12/29/2020 12:03:24",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> did you test if the improvement on validation translates to LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130962,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/29/2020 13:02:52",
          "content": "<p>I'm afraid that to do so I'll have to change a lot of my inference pipeline. Hence I didn't check that</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127259,
      "author_name": "nadare",
      "author_url": "",
      "post_date": "12/26/2020 11:27:47",
      "content": "<p>I use the history of lectures. However, I think that if you don't treat them carefully when using them with questions, you will damage the model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1118551": "I've been seeing quite varied opinions on usage of lectures in the models. So it'll be nice to know how everyone is faring with these and how much are people gaining (or losing) after addition of lectures.\n\nMy stats are:\nModel : Transfomer-esque (Val: 0.79248, LV: 0.792)\nLectures : Included\nGain/~~Loss~~ after addition: ~0.003\n\nNot sure if the gain was que to randomness or purely due to the addition of lectures as I made some more changes to the architecture and can't compare against the original baseline again.",
    "1118597": "abdurrafae I've tried to use lectures in transformer too like with `[Q1, Q2, Q3, L1, Q4, Q5, L2, Q5 ...]` with new encoding for `Ln` to not conflict with `Qn` but it did not provide any gain. How do you use lectures in sequence? Dedicated embedding?",
    "1118612": "I encoded both of them using the same layer after ensuring that all content_ids are different for questions and lectures.\n\nThis increases a bit of preprocessing in the pipeline at inference but simplifies the model.",
    "1118631": "Thanks for the answer. That's exactly what I've done, and what answer value did you encode for lecture? I made the prediction to be softmax with 3 classes but results were better without lectures.",
    "1118632": "I'm masking the prediction for lectures from my loss function. So model only learns to predict the answer for questions.",
    "1118635": "What's your model size and current score for transformer? (I assume your team's score is from LGBM)\n\nI have 128 d_model and 4 num_layers. Not sure how much help would scaling up the d_model parameter would achieve and I don't have the luxury (GPU resources) for testing it out at the moment. Keeping this change for the last week so I can iterate the current sized model a bit faster.",
    "1118645": "d_model=256 and num_layers=4. We don't have any issue with GPU resources, limitation is more with host RAM. You get OOM on GPU? or you don't have enough GPU quota to train with more data?",
    "1118650": "It takes around 20-24 hours to train one model once using Kaggle GPUs. Then I don't have much quota left for training a new one in the same week.",
    "1118653": "Maybe you could try optimize training (dataloader), how long is one epoch? It's 12 minutes for me.",
    "1118702": "My batch generator (tf.keras) is able to iterate through the whole dataset in less than a minute. So I know the issue is in the long training time of the model. 1 epoch is around 30 mins and I train upto 42 epochs for now.",
    "1118703": "I see, so your modified architecture should increase a lot the total number of model parameters, like in the way you combinate the embeddings?",
    "1118725": "The model size isn't that large ~4 - 5 M parameters overall. I think it may be due to my own implementation of attention module and masks.",
    "1119092": "Anyone knows why there are some questions and lectures with the same ID? And why, on the other hand, they can have different parts?",
    "1119097": "I think that's just an artifact of the data generating process and the number coincide by chance. I don't think there is any significance of a matching lecture_id and question_id.",
    "1119171": "May be, yeah... And the parts are the same/equivalent? I'm thinking about integrating lectures on my Riid-based model, but with 2% of appearance is it even worth it?",
    "1119193": "I use the lecture from the beginning, so not able to say which is better. I make sure it is not used for loss calculation, and not included in the inference during submitting.",
    "1119196": "I don't know which framework you use, @abdurrafae . I use Google Colab Pro + Tensorflow + TPU -> each epoch (~1.024 M sequences of window size 128) take 6m. Maybe it is late for you to setup TPU, but you should consider this (with Google Colab , pro or not) as an option.",
    "1119228": "I use TF keras. My dataset can be loaded on to RAM all at once so I think setting up the TPU shouldn't be that hard. I'll definitely try that out for my next model training.\n\nI also have almost same number of sequences btw :D\n\nIf you don't mind telling, how big is your model as in d_model and num_layers for the 6 min training time? and can you try it out on GPU for one epoch so I can benchmark off of it.",
    "1119239": "d_dim = 256, ffn_dim = 1024, n_layers = 4. Don't remember exactly the timing for GPU, but for each batch (128 sequences of window size 128), looks like 0.25 second, while with TPU, the same number of sequences is processed by TPU in 0.04 - 0.05 seconds.\n\nI use TFRecord files and there are quite a lot transformation - so it might not make the full power of TPUs.",
    "1119245": "Thanks. Even if it only take me to 15-20 mins, I still save a few hours in each iteration of the model then.",
    "1119252": "That's the entire reason I opened this thread. To get to know it's worth :D",
    "1119629": "abdurrafae I made another try with lectures but it did not provide any gain on CV. I might be doing something wrong. I made sure to encode lectures_id to be after questions_id and mask them in loss. The probability to get lectures in a random sequence is quite low, did you force the selection to have some lectures?",
    "1119678": "I create my training dataset by taking windows from the original sequences so I'm sure that lectures would be included. \n\nThe thing with Lectures is that the hassle to include them in the pipeline in inference is too much and since I faced a gain when I included them, I'm quite reluctant to roll back. However, I've not sure how much an actual gain they give and wanted to get a feel for that in this thread. \n\nIt's easier to see the gain/importance in LGBM. I wouldn't ask what your features related to lectures are (if any) in your LGBM model, but it'll be really helpful to know if they are important enough or not.",
    "1119687": "Easy to answer, all attempts to use lectures so far in LGBM failed.",
    "1119688": "that is sad",
    "1119692": "I woke up this morning with a new model trained with lectures. It did not improve.",
    "1119934": "Although we shouldn't be emotional, but I still feel a bit disappointed that lectures doesn't help - It means that some very appealing ideas and intuitions don't apply to models - and we need to try a lot of options to find out.\n\nActually, at once, I had a feeling to rework on my dataset pipeline and inference pipeline to remove lectures (because I didn't know if including it hurt my score) - but it requires quite a lot of my work, so I didn't change it.",
    "1120249": "Will it not work even if we discard lectures that the users didnt really spend enough time? I suppose thru some crude ways we can get the time lecture watched? But the data is so sparse that maybe it is not helping. Intuition-wise it definitely should. I saw one top10er saying he uses lectures..though maybe he has considered it by default in his original design itself\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107993",
    "1121385": "yihdarshieh Can you confirm if we need all datatypes to be converted to int32/64 or float32 for using the TPUs?",
    "1127125": "Update:\nI set up a pipeline for data without lectures on the same model that I had fitted data with lectures.\n\nModel fitted on **data containing lectures** had a better train/val auc score on each epoch. I stopped the test after only 10 epochs though to preserve my quota. (difference in val_auc was around ~.002 by the 10th epoch)",
    "1127133": "Great to know lectures help, otherwise the kids won't never study ... :)",
    "1127137": "Excellent, and they would use this thread as proof lectures are not worth. 😆",
    "1127140": "abdurrafae Sorry for asking in this thread with another question - I saw once, in the early stage of this competition, you tried to used masked LM loss to pretrain the encoder, and said there was a 0.005 gain.\n\nNow, with your improved model, does masked LM loss pretraining encoder still help?\n\nI tried it, it gives no gain (even a bit worse, but not much). And when I checked the accuracy of token (i.e `content_id` of questions / lectures) prediction, the best I can get is about 28% - 34%.\n\nI really wonder if MLM helps here.",
    "1127148": "I've only used binary-cross entropy loss with masks for lectures so far. I also remember someone talking about MLM but I didn't understand much about it",
    "1127259": "I use the history of lectures. However, I think that if you don't treat them carefully when using them with questions, you will damage the model.",
    "1130905": "abdurrafae did you test if the improvement on validation translates to LB?",
    "1130962": "I'm afraid that to do so I'll have to change a lot of my inference pipeline. Hence I didn't check that"
  },
  "source": "meta"
}