{
  "id": 206719,
  "title": "Retired Hurt - Final thoughts before moving on",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206719",
  "author_name": "",
  "post_date": "2020-12-26T08:12:39.140448500Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Most likely, it looks like Kaggle is not going to fix their bug as it affects only my userID (and they are not too bothered about single individuals) and therefore I may have to move on to other competitions - ‘retired hurt’ using cricket terminology though in this case the ‘hurt’ is not physical. Before moving on, wanted to share my last couple of thoughts:</p>\n<p>Leveraging lectures into the model: Have seen few discussions on this. One group claims that lectures dont help, the other group has been using lectures since the beginning - so they dont know whether it helps or not and a third group are still unclear whether to use them or not. The problem is one of sparsity. The lecture data is too sparse for the model to assign meaningful attention weights to it. One option could be to amplify the signal coming out…so basically multiply the attention weight in case the timestep is a lecture by a large value and hope that now it starts contributing in a more meaningful way to the predictions. The challenge here is that ‘interactions’ are still likely to be a better predictor than lectures. A typical flow could be a lecture - followed by dozens of questions. The student seeing the lecture might mean one thing - but whether she has understood what is being explained by responding correctly to the questions that follow is another thing altogether.  Lectures ‘may’ help in deducing whether the student is going to answer the immediate few questions following the lecture, but in the long run it will not help much..and the prediction will shift to be more dependent on the interactions following the lecture (instead of the lecture itself). The other reason why this happens is that memory retained after watching a lecture declines rapidly. So does it make more sense to (a) retain lectures only if they are done in the same or previous session and (b) coupled with giving it a much larger weight to amplify its impact. This combination ‘may’ work.</p>\n<p>Which brings us to the topic of forgetting. Haven’t seen any discussions or kernels around that. My initial thought was that Saint+ models forgetting but I dont think they do so explicitly. They give the elapsed time to the model and probably hope that the attention weights are learnt in such a way that more emphasis is placed on recent interactions. One option is to just ignore positional embedding and elapsed times and focus solely on building the attention weights with the interaction data alone. We dont need positional info at all. But we could build in a ‘forgetting’ module into the model based on the elapsed time. One crude way could be generate the attention weights and then give a higher weightage to the last few 20% transactions, a lower weightage to the first 20% and leave the rest untouched. AKT paper does some jazzy mathematical stuff and builds in an exponential decay - the exact details of which I couldnt understand but the best option would be to categorise the past sequence into 4 categories - same session, prev 24 hours, prev week and all the rest - and let the model itself learn the attention decay(or augmentation)  rates for these 4 categories. For e.g. (note this is just an example - I haven’t tested) we may end up with a augmentation of 1.5 times for the same session attention weights, 1.2 for prev 24 hours, 1 for prev week and 0.5 for rest of the timesteps. These values are multiplied to the attention weights generated.<br>\nNote - In that case we dont need positional info (if we are explicitly modelling forgetting). It is also much cleaner this way…since the initial attention weights are solely generated by only the past interactions…and final attention weights are generated by augmenting or decaying this with the learnt weights</p>",
  "messages": [
    {
      "id": "1127067",
      "postDate": "12/26/2020 08:12:39",
      "content": "<p>Most likely, it looks like Kaggle is not going to fix their bug as it affects only my userID (and they are not too bothered about single individuals) and therefore I may have to move on to other competitions - ‘retired hurt’ using cricket terminology though in this case the ‘hurt’ is not physical. Before moving on, wanted to share my last couple of thoughts:</p>\n<p>Leveraging lectures into the model: Have seen few discussions on this. One group claims that lectures dont help, the other group has been using lectures since the beginning - so they dont know whether it helps or not and a third group are still unclear whether to use them or not. The problem is one of sparsity. The lecture data is too sparse for the model to assign meaningful attention weights to it. One option could be to amplify the signal coming out…so basically multiply the attention weight in case the timestep is a lecture by a large value and hope that now it starts contributing in a more meaningful way to the predictions. The challenge here is that ‘interactions’ are still likely to be a better predictor than lectures. A typical flow could be a lecture - followed by dozens of questions. The student seeing the lecture might mean one thing - but whether she has understood what is being explained by responding correctly to the questions that follow is another thing altogether.  Lectures ‘may’ help in deducing whether the student is going to answer the immediate few questions following the lecture, but in the long run it will not help much..and the prediction will shift to be more dependent on the interactions following the lecture (instead of the lecture itself). The other reason why this happens is that memory retained after watching a lecture declines rapidly. So does it make more sense to (a) retain lectures only if they are done in the same or previous session and (b) coupled with giving it a much larger weight to amplify its impact. This combination ‘may’ work.</p>\n<p>Which brings us to the topic of forgetting. Haven’t seen any discussions or kernels around that. My initial thought was that Saint+ models forgetting but I dont think they do so explicitly. They give the elapsed time to the model and probably hope that the attention weights are learnt in such a way that more emphasis is placed on recent interactions. One option is to just ignore positional embedding and elapsed times and focus solely on building the attention weights with the interaction data alone. We dont need positional info at all. But we could build in a ‘forgetting’ module into the model based on the elapsed time. One crude way could be generate the attention weights and then give a higher weightage to the last few 20% transactions, a lower weightage to the first 20% and leave the rest untouched. AKT paper does some jazzy mathematical stuff and builds in an exponential decay - the exact details of which I couldnt understand but the best option would be to categorise the past sequence into 4 categories - same session, prev 24 hours, prev week and all the rest - and let the model itself learn the attention decay(or augmentation)  rates for these 4 categories. For e.g. (note this is just an example - I haven’t tested) we may end up with a augmentation of 1.5 times for the same session attention weights, 1.2 for prev 24 hours, 1 for prev week and 0.5 for rest of the timesteps. These values are multiplied to the attention weights generated.<br>\nNote - In that case we dont need positional info (if we are explicitly modelling forgetting). It is also much cleaner this way…since the initial attention weights are solely generated by only the past interactions…and final attention weights are generated by augmenting or decaying this with the learnt weights</p>",
      "rawMarkdown": "Most likely, it looks like Kaggle is not going to fix their bug as it affects only my userID (and they are not too bothered about single individuals) and therefore I may have to move on to other competitions - ‘retired hurt’ using cricket terminology though in this case the ‘hurt’ is not physical. Before moving on, wanted to share my last couple of thoughts:\n\nLeveraging lectures into the model: Have seen few discussions on this. One group claims that lectures dont help, the other group has been using lectures since the beginning - so they dont know whether it helps or not and a third group are still unclear whether to use them or not. The problem is one of sparsity. The lecture data is too sparse for the model to assign meaningful attention weights to it. One option could be to amplify the signal coming out…so basically multiply the attention weight in case the timestep is a lecture by a large value and hope that now it starts contributing in a more meaningful way to the predictions. The challenge here is that ‘interactions’ are still likely to be a better predictor than lectures. A typical flow could be a lecture - followed by dozens of questions. The student seeing the lecture might mean one thing - but whether she has understood what is being explained by responding correctly to the questions that follow is another thing altogether.  Lectures ‘may’ help in deducing whether the student is going to answer the immediate few questions following the lecture, but in the long run it will not help much..and the prediction will shift to be more dependent on the interactions following the lecture (instead of the lecture itself). The other reason why this happens is that memory retained after watching a lecture declines rapidly. So does it make more sense to (a) retain lectures only if they are done in the same or previous session and (b) coupled with giving it a much larger weight to amplify its impact. This combination ‘may’ work.\n\nWhich brings us to the topic of forgetting. Haven’t seen any discussions or kernels around that. My initial thought was that Saint+ models forgetting but I dont think they do so explicitly. They give the elapsed time to the model and probably hope that the attention weights are learnt in such a way that more emphasis is placed on recent interactions. One option is to just ignore positional embedding and elapsed times and focus solely on building the attention weights with the interaction data alone. We dont need positional info at all. But we could build in a ‘forgetting’ module into the model based on the elapsed time. One crude way could be generate the attention weights and then give a higher weightage to the last few 20% transactions, a lower weightage to the first 20% and leave the rest untouched. AKT paper does some jazzy mathematical stuff and builds in an exponential decay - the exact details of which I couldnt understand but the best option would be to categorise the past sequence into 4 categories - same session, prev 24 hours, prev week and all the rest - and let the model itself learn the attention decay(or augmentation)  rates for these 4 categories. For e.g. (note this is just an example - I haven’t tested) we may end up with a augmentation of 1.5 times for the same session attention weights, 1.2 for prev 24 hours, 1 for prev week and 0.5 for rest of the timesteps. These values are multiplied to the attention weights generated.\nNote - In that case we dont need positional info (if we are explicitly modelling forgetting). It is also much cleaner this way...since the initial attention weights are solely generated by only the past interactions...and final attention weights are generated by augmenting or decaying this with the learnt weights",
      "votes": null
    },
    {
      "id": "1127283",
      "postDate": "12/26/2020 11:58:29",
      "content": "<p>What is their bug, I think I am also affected by a bug (no GPU enabled for submission), and they are ignoring all of my email so far.</p>",
      "rawMarkdown": "What is their bug, I think I am also affected by a bug (no GPU enabled for submission), and they are ignoring all of my email so far.",
      "votes": null
    },
    {
      "id": "1127305",
      "postDate": "12/26/2020 12:27:12",
      "content": "<p>No <a href=\"https://www.kaggle.com/abdessalemboukil\" target=\"_blank\">@abdessalemboukil</a> mine is different. In my case, the submission executes but does not get counted or even displayed. Nothing related to GPU or CPU. </p>\n<p>I think I know now the root cause of this, but that is a different story. If you are facing the problem I am facing let me know</p>",
      "rawMarkdown": "No @abdessalemboukil mine is different. In my case, the submission executes but does not get counted or even displayed. Nothing related to GPU or CPU. \n\nI think I know now the root cause of this, but that is a different story. If you are facing the problem I am facing let me know",
      "votes": null
    },
    {
      "id": "1127316",
      "postDate": "12/26/2020 12:39:19",
      "content": "<p>That sucks man, no I am facing a different problem. But I think kaggle staff support is nearly non existing, apart of course from the bot that sends you the number of your ticket. Reading through older threads here confirms that.</p>",
      "rawMarkdown": "That sucks man, no I am facing a different problem. But I think kaggle staff support is nearly non existing, apart of course from the bot that sends you the number of your ticket. Reading through older threads here confirms that.",
      "votes": null
    },
    {
      "id": "1127323",
      "postDate": "12/26/2020 12:46:30",
      "content": "<p>I was viewing the support requests from the past and see that they used to respond better earlier (2-3 years back). But now they dont care about individual users any more. Possibly they have grown big now and hence this change. Unless we join voices together (the seniors in the forum need to support us) else they are not going to care. One thing I can tell is that this could happen to anybody in their next competition.</p>",
      "rawMarkdown": "I was viewing the support requests from the past and see that they used to respond better earlier (2-3 years back). But now they dont care about individual users any more. Possibly they have grown big now and hence this change. Unless we join voices together (the seniors in the forum need to support us) else they are not going to care. One thing I can tell is that this could happen to anybody in their next competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1127283,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "12/26/2020 11:58:29",
      "content": "<p>What is their bug, I think I am also affected by a bug (no GPU enabled for submission), and they are ignoring all of my email so far.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1127305,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "12/26/2020 12:27:12",
      "content": "<p>No <a href=\"https://www.kaggle.com/abdessalemboukil\" target=\"_blank\">@abdessalemboukil</a> mine is different. In my case, the submission executes but does not get counted or even displayed. Nothing related to GPU or CPU. </p>\n<p>I think I know now the root cause of this, but that is a different story. If you are facing the problem I am facing let me know</p>",
      "votes": null,
      "replies": [
        {
          "id": 1127316,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/26/2020 12:39:19",
          "content": "<p>That sucks man, no I am facing a different problem. But I think kaggle staff support is nearly non existing, apart of course from the bot that sends you the number of your ticket. Reading through older threads here confirms that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127323,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/26/2020 12:46:30",
          "content": "<p>I was viewing the support requests from the past and see that they used to respond better earlier (2-3 years back). But now they dont care about individual users any more. Possibly they have grown big now and hence this change. Unless we join voices together (the seniors in the forum need to support us) else they are not going to care. One thing I can tell is that this could happen to anybody in their next competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1127067": "Most likely, it looks like Kaggle is not going to fix their bug as it affects only my userID (and they are not too bothered about single individuals) and therefore I may have to move on to other competitions - ‘retired hurt’ using cricket terminology though in this case the ‘hurt’ is not physical. Before moving on, wanted to share my last couple of thoughts:\n\nLeveraging lectures into the model: Have seen few discussions on this. One group claims that lectures dont help, the other group has been using lectures since the beginning - so they dont know whether it helps or not and a third group are still unclear whether to use them or not. The problem is one of sparsity. The lecture data is too sparse for the model to assign meaningful attention weights to it. One option could be to amplify the signal coming out…so basically multiply the attention weight in case the timestep is a lecture by a large value and hope that now it starts contributing in a more meaningful way to the predictions. The challenge here is that ‘interactions’ are still likely to be a better predictor than lectures. A typical flow could be a lecture - followed by dozens of questions. The student seeing the lecture might mean one thing - but whether she has understood what is being explained by responding correctly to the questions that follow is another thing altogether.  Lectures ‘may’ help in deducing whether the student is going to answer the immediate few questions following the lecture, but in the long run it will not help much..and the prediction will shift to be more dependent on the interactions following the lecture (instead of the lecture itself). The other reason why this happens is that memory retained after watching a lecture declines rapidly. So does it make more sense to (a) retain lectures only if they are done in the same or previous session and (b) coupled with giving it a much larger weight to amplify its impact. This combination ‘may’ work.\n\nWhich brings us to the topic of forgetting. Haven’t seen any discussions or kernels around that. My initial thought was that Saint+ models forgetting but I dont think they do so explicitly. They give the elapsed time to the model and probably hope that the attention weights are learnt in such a way that more emphasis is placed on recent interactions. One option is to just ignore positional embedding and elapsed times and focus solely on building the attention weights with the interaction data alone. We dont need positional info at all. But we could build in a ‘forgetting’ module into the model based on the elapsed time. One crude way could be generate the attention weights and then give a higher weightage to the last few 20% transactions, a lower weightage to the first 20% and leave the rest untouched. AKT paper does some jazzy mathematical stuff and builds in an exponential decay - the exact details of which I couldnt understand but the best option would be to categorise the past sequence into 4 categories - same session, prev 24 hours, prev week and all the rest - and let the model itself learn the attention decay(or augmentation)  rates for these 4 categories. For e.g. (note this is just an example - I haven’t tested) we may end up with a augmentation of 1.5 times for the same session attention weights, 1.2 for prev 24 hours, 1 for prev week and 0.5 for rest of the timesteps. These values are multiplied to the attention weights generated.\nNote - In that case we dont need positional info (if we are explicitly modelling forgetting). It is also much cleaner this way...since the initial attention weights are solely generated by only the past interactions...and final attention weights are generated by augmenting or decaying this with the learnt weights",
    "1127283": "What is their bug, I think I am also affected by a bug (no GPU enabled for submission), and they are ignoring all of my email so far.",
    "1127305": "No @abdessalemboukil mine is different. In my case, the submission executes but does not get counted or even displayed. Nothing related to GPU or CPU. \n\nI think I know now the root cause of this, but that is a different story. If you are facing the problem I am facing let me know",
    "1127316": "That sucks man, no I am facing a different problem. But I think kaggle staff support is nearly non existing, apart of course from the bot that sends you the number of your ticket. Reading through older threads here confirms that.",
    "1127323": "I was viewing the support requests from the past and see that they used to respond better earlier (2-3 years back). But now they dont care about individual users any more. Possibly they have grown big now and hence this change. Unless we join voices together (the seniors in the forum need to support us) else they are not going to care. One thing I can tell is that this could happen to anybody in their next competition."
  },
  "source": "meta"
}