{
  "id": 205368,
  "title": "Look evenly to all data",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205368",
  "author_name": "Rodolphe Lampe",
  "post_date": "2020-12-19T20:52:16.004000",
  "votes": 31,
  "comment_count": 45,
  "views": 0,
  "content": "<p>In most notebooks on Transformers architecture, we loop over students and take 100 random consecutive interactions. Its largely suboptimal and I'll explain why and how to fix it.</p>\n<h2>The problem</h2>\n<p>Suppose you have a student1 with 100 rows and student2 with 3000 rows. You'll take the 100 rows of student1 at each epoch but you'll take only 100 rows from student2 at each epochs so each rows of students is seen 30 times less often than those of student1, leading to very quick overfitting (at least for me). My transformers was overfitting in less than 10 epochs.</p>\n<h2>A first bad solution</h2>\n<p>We could use weights in the loss with coefficients that balance the effect describe. However, there is a lot of noise in the sampling of rows so that, for student having many rows, it's very like that over a few dozen of epochs you won't see the data. The random choosing of interactions is not appropriate.</p>\n<h2>Solution 1</h2>\n<p>Take each 100 rows of each students. For example for student 2, you take rows 0-99, then 100-199, etc.<br>\nThis extremly simple thing improves a lot the overfitting problem and the score.</p>\n<h2>Improvement 1</h2>\n<p>The row 100 for example is seen only in a context where it's the first row and your model won't learn the interaction it has with rows 1-99. To avoid that, you'd like to make overlapping sliding windows. But again, you'll face the problem that some rows are seen more often than not : for example if you do it with 150 rows: you see 0-99 and then 50-149 and you've seen 2 times more often rows 50-99</p>\n<h2>Best solution I found</h2>\n<p>Use overlapping sliding windows while computing weights such that <strong>each row has exactly a weight one over an entire epoch</strong>. The maths are a bit complex so I'll give the formula :</p>\n<pre><code>def get_list_of_positions_of_one_sequence(length: int, max_length: int, overlap: int):\n    \"\"\"Return the start position and the associated weight\"\"\"\n    from math import ceil, floor\n    weights = list()\n    for i in range(length):\n        weight = 1 + min(floor(i/overlap), max(0, ceil((length-max_length)/overlap))) -\\\n                 max(0, ceil((i - (max_length - 1)) / overlap))\n        weights.append(weight)\n    weights = 1/np.array(weights)\n    start = 0\n    while True:\n        yield start, weights[start:start+max_length]\n        if start+max_length&gt;=length:\n            break\n        start += overlap\n</code></pre>\n<p>If you have length interactions and using max_length = 100 and an overlap of 50 (you translate of 50 interactions every time), the get_list_of_positions_of_one_sequence(length, 100, 50) will yield the start position and the associated 100 weights you can use.</p>",
  "messages": [
    {
      "id": 1119240,
      "postDate": "2020-12-19T20:52:16.003Z",
      "content": "<p>In most notebooks on Transformers architecture, we loop over students and take 100 random consecutive interactions. Its largely suboptimal and I'll explain why and how to fix it.</p>\n<h2>The problem</h2>\n<p>Suppose you have a student1 with 100 rows and student2 with 3000 rows. You'll take the 100 rows of student1 at each epoch but you'll take only 100 rows from student2 at each epochs so each rows of students is seen 30 times less often than those of student1, leading to very quick overfitting (at least for me). My transformers was overfitting in less than 10 epochs.</p>\n<h2>A first bad solution</h2>\n<p>We could use weights in the loss with coefficients that balance the effect describe. However, there is a lot of noise in the sampling of rows so that, for student having many rows, it's very like that over a few dozen of epochs you won't see the data. The random choosing of interactions is not appropriate.</p>\n<h2>Solution 1</h2>\n<p>Take each 100 rows of each students. For example for student 2, you take rows 0-99, then 100-199, etc.<br>\nThis extremly simple thing improves a lot the overfitting problem and the score.</p>\n<h2>Improvement 1</h2>\n<p>The row 100 for example is seen only in a context where it's the first row and your model won't learn the interaction it has with rows 1-99. To avoid that, you'd like to make overlapping sliding windows. But again, you'll face the problem that some rows are seen more often than not : for example if you do it with 150 rows: you see 0-99 and then 50-149 and you've seen 2 times more often rows 50-99</p>\n<h2>Best solution I found</h2>\n<p>Use overlapping sliding windows while computing weights such that <strong>each row has exactly a weight one over an entire epoch</strong>. The maths are a bit complex so I'll give the formula :</p>\n<pre><code>def get_list_of_positions_of_one_sequence(length: int, max_length: int, overlap: int):\n    \"\"\"Return the start position and the associated weight\"\"\"\n    from math import ceil, floor\n    weights = list()\n    for i in range(length):\n        weight = 1 + min(floor(i/overlap), max(0, ceil((length-max_length)/overlap))) -\\\n                 max(0, ceil((i - (max_length - 1)) / overlap))\n        weights.append(weight)\n    weights = 1/np.array(weights)\n    start = 0\n    while True:\n        yield start, weights[start:start+max_length]\n        if start+max_length&gt;=length:\n            break\n        start += overlap\n</code></pre>\n<p>If you have length interactions and using max_length = 100 and an overlap of 50 (you translate of 50 interactions every time), the get_list_of_positions_of_one_sequence(length, 100, 50) will yield the start position and the associated 100 weights you can use.</p>",
      "rawMarkdown": "In most notebooks on Transformers architecture, we loop over students and take 100 random consecutive interactions. Its largely suboptimal and I'll explain why and how to fix it.\n\n## The problem\n\nSuppose you have a student1 with 100 rows and student2 with 3000 rows. You'll take the 100 rows of student1 at each epoch but you'll take only 100 rows from student2 at each epochs so each rows of students is seen 30 times less often than those of student1, leading to very quick overfitting (at least for me). My transformers was overfitting in less than 10 epochs.\n\n## A first bad solution\n\nWe could use weights in the loss with coefficients that balance the effect describe. However, there is a lot of noise in the sampling of rows so that, for student having many rows, it's very like that over a few dozen of epochs you won't see the data. The random choosing of interactions is not appropriate.\n\n## Solution 1\n\nTake each 100 rows of each students. For example for student 2, you take rows 0-99, then 100-199, etc.\nThis extremly simple thing improves a lot the overfitting problem and the score.\n\n## Improvement 1\n\nThe row 100 for example is seen only in a context where it's the first row and your model won't learn the interaction it has with rows 1-99. To avoid that, you'd like to make overlapping sliding windows. But again, you'll face the problem that some rows are seen more often than not : for example if you do it with 150 rows: you see 0-99 and then 50-149 and you've seen 2 times more often rows 50-99\n\n## Best solution I found\n\nUse overlapping sliding windows while computing weights such that **each row has exactly a weight one over an entire epoch**. The maths are a bit complex so I'll give the formula :\n\n```python\ndef get_list_of_positions_of_one_sequence(length: int, max_length: int, overlap: int):\n    \"\"\"Return the start position and the associated weight\"\"\"\n    from math import ceil, floor\n    weights = list()\n    for i in range(length):\n        weight = 1 + min(floor(i/overlap), max(0, ceil((length-max_length)/overlap))) -\\\n                 max(0, ceil((i - (max_length - 1)) / overlap))\n        weights.append(weight)\n    weights = 1/np.array(weights)\n    start = 0\n    while True:\n        yield start, weights[start:start+max_length]\n        if start+max_length>=length:\n            break\n        start += overlap\n\n```\n\nIf you have length interactions and using max_length = 100 and an overlap of 50 (you translate of 50 interactions every time), the get_list_of_positions_of_one_sequence(length, 100, 50) will yield the start position and the associated 100 weights you can use.",
      "votes": 30
    },
    {
      "id": 1119256,
      "postDate": "2020-12-19T21:03:13.633Z",
      "content": "<p>Sampling strategy:</p>\n<ul>\n<li>Take all user ids in training data.</li>\n<li>Compute their lengths and cap the result to 500.</li>\n<li>Scale the prior to make them sum 1 (probability distribution).</li>\n<li>Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.</li>\n<li>Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.</li>\n<li>Repeat each epoch.</li>\n</ul>\n<p>Why?</p>\n<ul>\n<li>If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times.</li>\n<li>Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.</li>\n</ul>\n<p>(Extensive discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\" target=\"_blank\">here</a>)</p>",
      "rawMarkdown": "Sampling strategy:\n\n- Take all user ids in training data.\n- Compute their lengths and cap the result to 500.\n- Scale the prior to make them sum 1 (probability distribution).\n- Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.\n- Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.\n- Repeat each epoch.\n\nWhy?\n\n- If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times.\n- Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.\n\n(Extensive discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632))",
      "votes": 13,
      "replies": [
        {
          "id": 1119272,
          "postDate": "2020-12-19T21:23:07.813Z",
          "content": "<p>Thanks ! I missed the discussion, I'll read it.<br>\nAbout avoiding super prolific users, why should we do that ? There could be also prolific users in the test set ?!</p>",
          "rawMarkdown": "Thanks ! I missed the discussion, I'll read it.\nAbout avoiding super prolific users, why should we do that ? There could be also prolific users in the test set ?!"
        },
        {
          "id": 1119282,
          "postDate": "2020-12-19T21:39:50.323Z",
          "content": "<p>Because their sequences would be appearing too many times since they would have an excessive sampling probability value.</p>",
          "rawMarkdown": "Because their sequences would be appearing too many times since they would have an excessive sampling probability value."
        },
        {
          "id": 1119283,
          "postDate": "2020-12-19T21:41:19.263Z",
          "content": "<p>Also using <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> approach that works. Thanks for additional solutions.</p>",
          "rawMarkdown": "Also using @claverru approach that works. Thanks for additional solutions."
        },
        {
          "id": 1119285,
          "postDate": "2020-12-19T21:43:18.293Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, not exactly. For a user with long history, say 5000, it contains much more subsequences, and they are different even they belongs to the same user.</p>\n<p>I am not saying your approach is wrong, just saying you above argument is not very strong for explaining your choice. </p>\n<p>Who knows, seeing (truncated) sequences from the same users (even they are different) might be harmful as you point</p>",
          "rawMarkdown": "@claverru, not exactly. For a user with long history, say 5000, it contains much more subsequences, and they are different even they belongs to the same user.\n\nI am not saying your approach is wrong, just saying you above argument is not very strong for explaining your choice. \n\nWho knows, seeing (truncated) sequences from the same users (even they are different) might be harmful as you point"
        },
        {
          "id": 1119856,
          "postDate": "2020-12-20T12:06:15.007Z",
          "content": "<p>Why should we balance the number of samples by user ? The metric is AUC so over rows and not a metric averaged over users. So we should indeed look at all users in an unbalanced way (more often users with more rows) no ?</p>",
          "rawMarkdown": "Why should we balance the number of samples by user ? The metric is AUC so over rows and not a metric averaged over users. So we should indeed look at all users in an unbalanced way (more often users with more rows) no ?",
          "votes": 2
        }
      ]
    },
    {
      "id": 1119255,
      "postDate": "2020-12-19T21:03:11.633Z",
      "content": "<p>Thank you for sharing, that's really nice. I want to point out that, <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>  and me also have some similar approach to this issue <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1116239\" target=\"_blank\">here</a> - despite some differences.</p>",
      "rawMarkdown": "Thank you for sharing, that's really nice. I want to point out that, @claverru  and me also have some similar approach to this issue [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1116239) - despite some differences.",
      "votes": 3
    },
    {
      "id": 1119250,
      "postDate": "2020-12-19T20:58:55.440Z",
      "content": "<p>I'm at the improvement I level at the moment. </p>\n<p>Best solution sound interesting, not sure how much of a gain would that bring. Can you share score of before and after for motivation :D</p>",
      "rawMarkdown": "I'm at the improvement I level at the moment. \n\nBest solution sound interesting, not sure how much of a gain would that bring. Can you share score of before and after for motivation :D",
      "votes": 1,
      "replies": [
        {
          "id": 1119254,
          "postDate": "2020-12-19T21:02:47.450Z",
          "content": "<p>Sorry, I didn't saved the intermediate results for that step. If anyone tries it, I'm interested too to have clear quantitative results of the final step.<br>\nAnyway, I don't think you should prioritize it over other important stuffs but if at some point you've no more ideas to try, it might be interesting to test it. The most important part was the improvement I level.</p>",
          "rawMarkdown": "Sorry, I didn't saved the intermediate results for that step. If anyone tries it, I'm interested too to have clear quantitative results of the final step.\nAnyway, I don't think you should prioritize it over other important stuffs but if at some point you've no more ideas to try, it might be interesting to test it. The most important part was the improvement I level."
        }
      ]
    },
    {
      "id": 1127340,
      "postDate": "2020-12-26T13:04:22.587Z",
      "content": "<p>Brilliant. Deserves more upvotes. One quick question though. What is wrong with the strategy of breaking a user&gt;100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?</p>\n<p>of course claveru point is also is much valid. There are many starter sequences (student dropouts) and we would need to find out simple ways to overcome that.</p>",
      "rawMarkdown": "Brilliant. Deserves more upvotes. One quick question though. What is wrong with the strategy of breaking a user>100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?\n\nof course claveru point is also is much valid. There are many starter sequences (student dropouts) and we would need to find out simple ways to overcome that.",
      "replies": [
        {
          "id": 1127345,
          "postDate": "2020-12-26T13:07:18.877Z",
          "content": "<blockquote>\n  <p>What is wrong with the strategy of breaking a user&gt;100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?</p>\n</blockquote>\n<p>This looses the information that the next (samples in the) batch of 100 seq's is anyway related to the previous. So we add the last 50 (let's say) to the next window's start as it's first 50 items. But i feel that if you will end up shuffling, then it's kinda that way pretty much nonetheless. </p>\n<p>I don't know whether my understanding is correct or not, so take it lightly.</p>\n<p>Edit-: It's a good idea to try and see, nonetheless!</p>",
          "rawMarkdown": ">What is wrong with the strategy of breaking a user>100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?\n\n\nThis looses the information that the next (samples in the) batch of 100 seq's is anyway related to the previous. So we add the last 50 (let's say) to the next window's start as it's first 50 items. But i feel that if you will end up shuffling, then it's kinda that way pretty much nonetheless. \n\nI don't know whether my understanding is correct or not, so take it lightly.\n\nEdit-: It's a good idea to try and see, nonetheless!",
          "votes": 1
        },
        {
          "id": 1127352,
          "postDate": "2020-12-26T13:10:33.670Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> your kernels on Pytorch have been very helpful to me and also your various discussions. I would always pay heed to your thoughts :)</p>",
          "rawMarkdown": "@adityaecdrid your kernels on Pytorch have been very helpful to me and also your various discussions. I would always pay heed to your thoughts :)"
        },
        {
          "id": 1127356,
          "postDate": "2020-12-26T13:13:58.007Z",
          "content": "<p>Just thinking whether the added processing would make a difference to the scores. Had there been a shortage of data we should extract every ounce of info from them. But I think here given the volumes of data, we may not lose much. But of course the algorithm suggested is a more thorough one for sure! definitely a useful suggestion</p>",
          "rawMarkdown": "Just thinking whether the added processing would make a difference to the scores. Had there been a shortage of data we should extract every ounce of info from them. But I think here given the volumes of data, we may not lose much. But of course the algorithm suggested is a more thorough one for sure! definitely a useful suggestion"
        }
      ]
    },
    {
      "id": 1120689,
      "postDate": "2020-12-21T03:56:31.857Z",
      "content": "<p>Just curious, wouldn't it be possible for the model to learn those common sequences as well wich overlap b/w the sub-sequences from the same user?</p>",
      "rawMarkdown": "Just curious, wouldn't it be possible for the model to learn those common sequences as well wich overlap b/w the sub-sequences from the same user?"
    },
    {
      "id": 1119967,
      "postDate": "2020-12-20T13:48:16.093Z",
      "content": "<p><a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> - why your LB rank is not shown? You have no submission so far?</p>",
      "rawMarkdown": "@rodolphelampe - why your LB rank is not shown? You have no submission so far?",
      "replies": [
        {
          "id": 1120486,
          "postDate": "2020-12-20T20:58:21.787Z",
          "content": "<p>Yes, not yet. But I hope to have one maybe tommorow, I work on my IDE and I'm not used to the kernel submission and we had some memory issues that we just solved.</p>",
          "rawMarkdown": "Yes, not yet. But I hope to have one maybe tommorow, I work on my IDE and I'm not used to the kernel submission and we had some memory issues that we just solved.",
          "votes": 1
        },
        {
          "id": 1120512,
          "postDate": "2020-12-20T21:29:09.167Z",
          "content": "<p>Good luck. But just as an suggestion (even for future competitions), it might be good to setup the pipeline for submission as soon as possible - because things might get very unexpected when the code running on submission with hidden dataset …..</p>",
          "rawMarkdown": "Good luck. But just as an suggestion (even for future competitions), it might be good to setup the pipeline for submission as soon as possible - because things might get very unexpected when the code running on submission with hidden dataset .....",
          "votes": 1
        },
        {
          "id": 1120518,
          "postDate": "2020-12-20T21:37:44.360Z",
          "content": "<p>I second this point. I actually took the whole first month setting up the pipeline. It's better to start with a simple one and iterate your ideas early on.</p>",
          "rawMarkdown": "I second this point. I actually took the whole first month setting up the pipeline. It's better to start with a simple one and iterate your ideas early on.",
          "votes": 1
        },
        {
          "id": 1120520,
          "postDate": "2020-12-20T21:40:31.587Z",
          "content": "<p>me too, the total time of setup and debug, definitely &gt; 1 month for me … </p>",
          "rawMarkdown": "me too, the total time of setup and debug, definitely > 1 month for me ... "
        },
        {
          "id": 1120522,
          "postDate": "2020-12-20T21:41:36.357Z",
          "content": "<p>Thanks for the advice, I see that now and I hope I will make it but yeah it takes a lot of time !</p>",
          "rawMarkdown": "Thanks for the advice, I see that now and I hope I will make it but yeah it takes a lot of time !",
          "votes": 1
        },
        {
          "id": 1120523,
          "postDate": "2020-12-20T21:42:44.177Z",
          "content": "<p>When I saw people have &gt;  0.79x LB with only a few submissions, I felt that I am so stupid (why other people are doing much better than me in much less time ….) </p>",
          "rawMarkdown": "When I saw people have >  0.79x LB with only a few submissions, I felt that I am so stupid (why other people are doing much better than me in much less time ....) "
        },
        {
          "id": 1135473,
          "postDate": "2021-01-02T09:26:33.710Z",
          "content": "<p>Your predictions were right, I've not yet managed to fix the errors I got when running my code on kaggle. I hope I'll fix it very soon !</p>",
          "rawMarkdown": "Your predictions were right, I've not yet managed to fix the errors I got when running my code on kaggle. I hope I'll fix it very soon !",
          "votes": 1
        },
        {
          "id": 1135492,
          "postDate": "2021-01-02T09:47:02.350Z",
          "content": "<p>This is the only place I predict well 😂 I am also struggling to finalize my pipeline for now</p>",
          "rawMarkdown": "This is the only place I predict well 😂 I am also struggling to finalize my pipeline for now",
          "votes": 1
        },
        {
          "id": 1142301,
          "postDate": "2021-01-07T09:55:07.263Z",
          "content": "<p>It took me 18 days 😬<br>\nI'm glad I could finish, it was so close </p>",
          "rawMarkdown": "It took me 18 days 😬\nI'm glad I could finish, it was so close ",
          "votes": 2
        },
        {
          "id": 1142304,
          "postDate": "2021-01-07T09:57:28.023Z",
          "content": "<p><a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> Congratulations on finally joining the LB :)</p>\n<p>It's really nice when everything comes together in the end.</p>",
          "rawMarkdown": "@rodolphelampe Congratulations on finally joining the LB :)\n\nIt's really nice when everything comes together in the end."
        },
        {
          "id": 1142338,
          "postDate": "2021-01-07T10:22:57.807Z",
          "content": "<p>Thanks and congratulations too, you're going to see a huge boost in your kaggle points / ranking, it's great how high in the LB you reached (still to be confirmed when the competitions ends but I doubt it will move much).</p>",
          "rawMarkdown": "Thanks and congratulations too, you're going to see a huge boost in your kaggle points / ranking, it's great how high in the LB you reached (still to be confirmed when the competitions ends but I doubt it will move much).",
          "votes": 1
        },
        {
          "id": 1142366,
          "postDate": "2021-01-07T10:37:20.933Z",
          "content": "<p>Yeah, this is the first competition I participated with a serious attitude and proper model inference pipeline.</p>\n<p>I'm glad that I've made it this far. I had an internal goal of reaching 0.81, so I'm quite happy with my work here.</p>\n<p>Although I was praying for the Green/Gold Zone 😅</p>",
          "rawMarkdown": "Yeah, this is the first competition I participated with a serious attitude and proper model inference pipeline.\n\nI'm glad that I've made it this far. I had an internal goal of reaching 0.81, so I'm quite happy with my work here.\n\nAlthough I was praying for the Green/Gold Zone 😅",
          "votes": 1
        },
        {
          "id": 1142380,
          "postDate": "2021-01-07T10:47:11.227Z",
          "content": "<p>Great for both of you, congratulations! I made quite improvement in modelling, CV 0.813, but can't get it working on submitting, a bit sad. It uses the per user  per question  historical statical information. Some bug in the local and submitting pipeline though😓</p>",
          "rawMarkdown": "Great for both of you, congratulations! I made quite improvement in modelling, CV 0.813, but can't get it working on submitting, a bit sad. It uses the per user  per question  historical statical information. Some bug in the local and submitting pipeline though😓"
        },
        {
          "id": 1142382,
          "postDate": "2021-01-07T10:48:44.700Z",
          "content": "<p>We still have ~12 hour left. Unless you've exhausted today's submission that is 😅</p>",
          "rawMarkdown": "We still have ~12 hour left. Unless you've exhausted today's submission that is 😅"
        },
        {
          "id": 1142384,
          "postDate": "2021-01-07T10:49:34.033Z",
          "content": "<p>How useful would it be to overfit the transformers? :)</p>\n<p>I've run out of all ideas and have 5 submissions left. Let's check it out </p>",
          "rawMarkdown": "How useful would it be to overfit the transformers? :)\n\nI've run out of all ideas and have 5 submissions left. Let's check it out ",
          "votes": 1
        },
        {
          "id": 1142388,
          "postDate": "2021-01-07T10:53:21.780Z",
          "content": "<p>No, i am still trying, don't want to give up. But considering the submission need to be finished before deadline, the time to submit remains just , say 5 hours…</p>",
          "rawMarkdown": "No, i am still trying, don't want to give up. But considering the submission need to be finished before deadline, the time to submit remains just , say 5 hours..."
        },
        {
          "id": 1142390,
          "postDate": "2021-01-07T10:55:42.273Z",
          "content": "<p>Best of luck mate! do you get the error immediately after submission or after some time?</p>",
          "rawMarkdown": "Best of luck mate! do you get the error immediately after submission or after some time?"
        },
        {
          "id": 1142393,
          "postDate": "2021-01-07T10:57:09.807Z",
          "content": "<p>The score is worse on LB , but on CV, i got from 0.808 to 0.813… input inconsistency I think</p>",
          "rawMarkdown": "The score is worse on LB , but on CV, i got from 0.808 to 0.813... input inconsistency I think"
        },
        {
          "id": 1142394,
          "postDate": "2021-01-07T10:58:42.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  For a potential next competition, I would like to make a team with you in an early stage, if you are also interesting?</p>",
          "rawMarkdown": "@abdurrafae  For a potential next competition, I would like to make a team with you in an early stage, if you are also interesting?"
        },
        {
          "id": 1142395,
          "postDate": "2021-01-07T10:58:59.023Z",
          "content": "<p>Could be some leakage as well. CV has been really well co-related to Public LB otherwise.</p>\n<p>Any idea how well CV is compared to Private LB?</p>",
          "rawMarkdown": "Could be some leakage as well. CV has been really well co-related to Public LB otherwise.\n\nAny idea how well CV is compared to Private LB?"
        },
        {
          "id": 1142404,
          "postDate": "2021-01-07T11:05:21.703Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I'm always interested in interesting competitions :)</p>\n<p>I would love to team up with you. I believe if we teamed up earlier (and you taught me how to use a TPU properly) we be in the gold for sure. 😏</p>\n<p>Spent almost like 2 weeks debugging different stuff in the TPU training pipeline but I can say it was worth it. Model trains in like 4-6 hours now max with ~20-30M parameters.</p>\n<p>Implementing a pseudorandom subsequence sampler was a headache and took 2 days but provided with a descent 0.003 LB improvement. (my most recent jump) <br>\nThe sampler still has a few bugs but I don't want to break it with just few hours of TPU quota left.</p>",
          "rawMarkdown": "@yihdarshieh I'm always interested in interesting competitions :)\n\nI would love to team up with you. I believe if we teamed up earlier (and you taught me how to use a TPU properly) we be in the gold for sure. 😏\n\nSpent almost like 2 weeks debugging different stuff in the TPU training pipeline but I can say it was worth it. Model trains in like 4-6 hours now max with ~20-30M parameters.\n\nImplementing a pseudorandom subsequence sampler was a headache and took 2 days but provided with a descent 0.003 LB improvement. (my most recent jump) \nThe sampler still has a few bugs but I don't want to break it with just few hours of TPU quota left.",
          "votes": 1
        },
        {
          "id": 1142405,
          "postDate": "2021-01-07T11:06:01.537Z",
          "content": "<p>I had same experience with you using same feature generating method and had a CV score of 0.816, but…</p>",
          "rawMarkdown": "I had same experience with you using same feature generating method and had a CV score of 0.816, but...",
          "votes": 1
        },
        {
          "id": 1142426,
          "postDate": "2021-01-07T11:16:12.620Z",
          "content": "<p>Looking forward to work with you in a next competition, I would need a few break before it😂</p>",
          "rawMarkdown": "Looking forward to work with you in a next competition, I would need a few break before it😂"
        },
        {
          "id": 1142427,
          "postDate": "2021-01-07T11:17:44.813Z",
          "content": "<p>Same here. I've been obsessed with the competition in last 2-3 weeks as I had holidays as well. :D</p>",
          "rawMarkdown": "Same here. I've been obsessed with the competition in last 2-3 weeks as I had holidays as well. :D",
          "votes": 2
        },
        {
          "id": 1142430,
          "postDate": "2021-01-07T11:18:39.510Z",
          "content": "<p>So mystery😨</p>",
          "rawMarkdown": "So mystery😨"
        },
        {
          "id": 1142440,
          "postDate": "2021-01-07T11:29:59.873Z",
          "content": "<p>I have 3 weeks vacation - and all of them but 3 days left went for Riiid ….</p>",
          "rawMarkdown": "I have 3 weeks vacation - and all of them but 3 days left went for Riiid ....",
          "votes": 1
        },
        {
          "id": 1142449,
          "postDate": "2021-01-07T11:38:31.647Z",
          "content": "<p>Ah, then I believe our situation to be quite the same.<br>\nI've submitted 2 overfitted transformer submissions. let's see what happens now. 😁</p>\n<p>Still have 4 hours TPU left. Thinking of trying again a new architecture that failed yesterday. 🤔</p>",
          "rawMarkdown": "Ah, then I believe our situation to be quite the same.\nI've submitted 2 overfitted transformer submissions. let's see what happens now. 😁\n\nStill have 4 hours TPU left. Thinking of trying again a new architecture that failed yesterday. 🤔",
          "votes": 1
        },
        {
          "id": 1143033,
          "postDate": "2021-01-07T18:16:18.277Z",
          "content": "<p>I found a bug (stupid one - variable naming …), I submitted 5 kernels. I  don't know if they can finish running before the deadline. Even if yes, I don't know what the results will be.</p>\n<p>Anyway, the competition is done, just waiting the results. Lessons learned, hope next time I can manage things better.</p>\n<p>Good luck, guys!</p>",
          "rawMarkdown": "I found a bug (stupid one - variable naming ...), I submitted 5 kernels. I  don't know if they can finish running before the deadline. Even if yes, I don't know what the results will be.\n\nAnyway, the competition is done, just waiting the results. Lessons learned, hope next time I can manage things better.\n\nGood luck, guys!"
        },
        {
          "id": 1144663,
          "postDate": "2021-01-08T15:46:05.920Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, do you have some insight on the failure of including user statistical performance? For me, I continued to found bugs in the input processing for submission, and I am able to get 0.802 private LB (0.799 on public LB), submitted after the deadline, but I think there is some bug in my submission pipeline. Because my submission pipeline give CV 0.807 while I got 0.813 CV when it is run with tensorflow dataset pipeline.</p>",
          "rawMarkdown": "@lihaorocky, do you have some insight on the failure of including user statistical performance? For me, I continued to found bugs in the input processing for submission, and I am able to get 0.802 private LB (0.799 on public LB), submitted after the deadline, but I think there is some bug in my submission pipeline. Because my submission pipeline give CV 0.807 while I got 0.813 CV when it is run with tensorflow dataset pipeline."
        }
      ]
    },
    {
      "id": 1119279,
      "postDate": "2020-12-19T21:33:25.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1119256,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-12-19T21:03:13.633000",
      "content": "<p>Sampling strategy:</p>\n<ul>\n<li>Take all user ids in training data.</li>\n<li>Compute their lengths and cap the result to 500.</li>\n<li>Scale the prior to make them sum 1 (probability distribution).</li>\n<li>Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.</li>\n<li>Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.</li>\n<li>Repeat each epoch.</li>\n</ul>\n<p>Why?</p>\n<ul>\n<li>If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times.</li>\n<li>Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.</li>\n</ul>\n<p>(Extensive discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\" target=\"_blank\">here</a>)</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1119272,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-19T21:23:07.813000",
          "content": "<p>Thanks ! I missed the discussion, I'll read it.<br>\nAbout avoiding super prolific users, why should we do that ? There could be also prolific users in the test set ?!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119282,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-19T21:39:50.323000",
          "content": "<p>Because their sequences would be appearing too many times since they would have an excessive sampling probability value.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119283,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-19T21:41:19.263000",
          "content": "<p>Also using <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> approach that works. Thanks for additional solutions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119285,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-19T21:43:18.293000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, not exactly. For a user with long history, say 5000, it contains much more subsequences, and they are different even they belongs to the same user.</p>\n<p>I am not saying your approach is wrong, just saying you above argument is not very strong for explaining your choice. </p>\n<p>Who knows, seeing (truncated) sequences from the same users (even they are different) might be harmful as you point</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119856,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-20T12:06:15.007000",
          "content": "<p>Why should we balance the number of samples by user ? The metric is AUC so over rows and not a metric averaged over users. So we should indeed look at all users in an unbalanced way (more often users with more rows) no ?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1119255,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-12-19T21:03:11.633000",
      "content": "<p>Thank you for sharing, that's really nice. I want to point out that, <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>  and me also have some similar approach to this issue <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1116239\" target=\"_blank\">here</a> - despite some differences.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1119250,
      "author_name": "AbdurRafae",
      "author_url": "",
      "post_date": "2020-12-19T20:58:55.440000",
      "content": "<p>I'm at the improvement I level at the moment. </p>\n<p>Best solution sound interesting, not sure how much of a gain would that bring. Can you share score of before and after for motivation :D</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1119254,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-19T21:02:47.450000",
          "content": "<p>Sorry, I didn't saved the intermediate results for that step. If anyone tries it, I'm interested too to have clear quantitative results of the final step.<br>\nAnyway, I don't think you should prioritize it over other important stuffs but if at some point you've no more ideas to try, it might be interesting to test it. The most important part was the improvement I level.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1127340,
      "author_name": "Allohvk",
      "author_url": "",
      "post_date": "2020-12-26T13:04:22.587000",
      "content": "<p>Brilliant. Deserves more upvotes. One quick question though. What is wrong with the strategy of breaking a user&gt;100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?</p>\n<p>of course claveru point is also is much valid. There are many starter sequences (student dropouts) and we would need to find out simple ways to overcome that.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1127345,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-26T13:07:18.877000",
          "content": "<blockquote>\n  <p>What is wrong with the strategy of breaking a user&gt;100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?</p>\n</blockquote>\n<p>This looses the information that the next (samples in the) batch of 100 seq's is anyway related to the previous. So we add the last 50 (let's say) to the next window's start as it's first 50 items. But i feel that if you will end up shuffling, then it's kinda that way pretty much nonetheless. </p>\n<p>I don't know whether my understanding is correct or not, so take it lightly.</p>\n<p>Edit-: It's a good idea to try and see, nonetheless!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1127352,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-26T13:10:33.670000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> your kernels on Pytorch have been very helpful to me and also your various discussions. I would always pay heed to your thoughts :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1127356,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-26T13:13:58.007000",
          "content": "<p>Just thinking whether the added processing would make a difference to the scores. Had there been a shortage of data we should extract every ounce of info from them. But I think here given the volumes of data, we may not lose much. But of course the algorithm suggested is a more thorough one for sure! definitely a useful suggestion</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1120689,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-12-21T03:56:31.857000",
      "content": "<p>Just curious, wouldn't it be possible for the model to learn those common sequences as well wich overlap b/w the sub-sequences from the same user?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1119967,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-12-20T13:48:16.093000",
      "content": "<p><a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> - why your LB rank is not shown? You have no submission so far?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1120486,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-20T20:58:21.787000",
          "content": "<p>Yes, not yet. But I hope to have one maybe tommorow, I work on my IDE and I'm not used to the kernel submission and we had some memory issues that we just solved.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120512,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-20T21:29:09.167000",
          "content": "<p>Good luck. But just as an suggestion (even for future competitions), it might be good to setup the pipeline for submission as soon as possible - because things might get very unexpected when the code running on submission with hidden dataset …..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120518,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-20T21:37:44.360000",
          "content": "<p>I second this point. I actually took the whole first month setting up the pipeline. It's better to start with a simple one and iterate your ideas early on.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120520,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-20T21:40:31.587000",
          "content": "<p>me too, the total time of setup and debug, definitely &gt; 1 month for me … </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1120522,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-20T21:41:36.357000",
          "content": "<p>Thanks for the advice, I see that now and I hope I will make it but yeah it takes a lot of time !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120523,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-20T21:42:44.177000",
          "content": "<p>When I saw people have &gt;  0.79x LB with only a few submissions, I felt that I am so stupid (why other people are doing much better than me in much less time ….) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135473,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2021-01-02T09:26:33.710000",
          "content": "<p>Your predictions were right, I've not yet managed to fix the errors I got when running my code on kaggle. I hope I'll fix it very soon !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1135492,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-02T09:47:02.350000",
          "content": "<p>This is the only place I predict well 😂 I am also struggling to finalize my pipeline for now</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142301,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2021-01-07T09:55:07.263000",
          "content": "<p>It took me 18 days 😬<br>\nI'm glad I could finish, it was so close </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1142304,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T09:57:28.023000",
          "content": "<p><a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> Congratulations on finally joining the LB :)</p>\n<p>It's really nice when everything comes together in the end.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142338,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2021-01-07T10:22:57.807000",
          "content": "<p>Thanks and congratulations too, you're going to see a huge boost in your kaggle points / ranking, it's great how high in the LB you reached (still to be confirmed when the competitions ends but I doubt it will move much).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142366,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T10:37:20.933000",
          "content": "<p>Yeah, this is the first competition I participated with a serious attitude and proper model inference pipeline.</p>\n<p>I'm glad that I've made it this far. I had an internal goal of reaching 0.81, so I'm quite happy with my work here.</p>\n<p>Although I was praying for the Green/Gold Zone 😅</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142380,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T10:47:11.227000",
          "content": "<p>Great for both of you, congratulations! I made quite improvement in modelling, CV 0.813, but can't get it working on submitting, a bit sad. It uses the per user  per question  historical statical information. Some bug in the local and submitting pipeline though😓</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142382,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T10:48:44.700000",
          "content": "<p>We still have ~12 hour left. Unless you've exhausted today's submission that is 😅</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142384,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T10:49:34.033000",
          "content": "<p>How useful would it be to overfit the transformers? :)</p>\n<p>I've run out of all ideas and have 5 submissions left. Let's check it out </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142388,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T10:53:21.780000",
          "content": "<p>No, i am still trying, don't want to give up. But considering the submission need to be finished before deadline, the time to submit remains just , say 5 hours…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142390,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T10:55:42.273000",
          "content": "<p>Best of luck mate! do you get the error immediately after submission or after some time?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142393,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T10:57:09.807000",
          "content": "<p>The score is worse on LB , but on CV, i got from 0.808 to 0.813… input inconsistency I think</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142394,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T10:58:42.963000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  For a potential next competition, I would like to make a team with you in an early stage, if you are also interesting?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142395,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T10:58:59.023000",
          "content": "<p>Could be some leakage as well. CV has been really well co-related to Public LB otherwise.</p>\n<p>Any idea how well CV is compared to Private LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142404,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T11:05:21.703000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I'm always interested in interesting competitions :)</p>\n<p>I would love to team up with you. I believe if we teamed up earlier (and you taught me how to use a TPU properly) we be in the gold for sure. 😏</p>\n<p>Spent almost like 2 weeks debugging different stuff in the TPU training pipeline but I can say it was worth it. Model trains in like 4-6 hours now max with ~20-30M parameters.</p>\n<p>Implementing a pseudorandom subsequence sampler was a headache and took 2 days but provided with a descent 0.003 LB improvement. (my most recent jump) <br>\nThe sampler still has a few bugs but I don't want to break it with just few hours of TPU quota left.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142405,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2021-01-07T11:06:01.537000",
          "content": "<p>I had same experience with you using same feature generating method and had a CV score of 0.816, but…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142426,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T11:16:12.620000",
          "content": "<p>Looking forward to work with you in a next competition, I would need a few break before it😂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142427,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T11:17:44.813000",
          "content": "<p>Same here. I've been obsessed with the competition in last 2-3 weeks as I had holidays as well. :D</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1142430,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T11:18:39.510000",
          "content": "<p>So mystery😨</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142440,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T11:29:59.873000",
          "content": "<p>I have 3 weeks vacation - and all of them but 3 days left went for Riiid ….</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142449,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T11:38:31.647000",
          "content": "<p>Ah, then I believe our situation to be quite the same.<br>\nI've submitted 2 overfitted transformer submissions. let's see what happens now. 😁</p>\n<p>Still have 4 hours TPU left. Thinking of trying again a new architecture that failed yesterday. 🤔</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1143033,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-07T18:16:18.277000",
          "content": "<p>I found a bug (stupid one - variable naming …), I submitted 5 kernels. I  don't know if they can finish running before the deadline. Even if yes, I don't know what the results will be.</p>\n<p>Anyway, the competition is done, just waiting the results. Lessons learned, hope next time I can manage things better.</p>\n<p>Good luck, guys!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144663,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-08T15:46:05.920000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, do you have some insight on the failure of including user statistical performance? For me, I continued to found bugs in the input processing for submission, and I am able to get 0.802 private LB (0.799 on public LB), submitted after the deadline, but I think there is some bug in my submission pipeline. Because my submission pipeline give CV 0.807 while I got 0.813 CV when it is run with tensorflow dataset pipeline.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1119279,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-19T21:33:25.650000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1119240": "In most notebooks on Transformers architecture, we loop over students and take 100 random consecutive interactions. Its largely suboptimal and I'll explain why and how to fix it.\n\n## The problem\n\nSuppose you have a student1 with 100 rows and student2 with 3000 rows. You'll take the 100 rows of student1 at each epoch but you'll take only 100 rows from student2 at each epochs so each rows of students is seen 30 times less often than those of student1, leading to very quick overfitting (at least for me). My transformers was overfitting in less than 10 epochs.\n\n## A first bad solution\n\nWe could use weights in the loss with coefficients that balance the effect describe. However, there is a lot of noise in the sampling of rows so that, for student having many rows, it's very like that over a few dozen of epochs you won't see the data. The random choosing of interactions is not appropriate.\n\n## Solution 1\n\nTake each 100 rows of each students. For example for student 2, you take rows 0-99, then 100-199, etc.\nThis extremly simple thing improves a lot the overfitting problem and the score.\n\n## Improvement 1\n\nThe row 100 for example is seen only in a context where it's the first row and your model won't learn the interaction it has with rows 1-99. To avoid that, you'd like to make overlapping sliding windows. But again, you'll face the problem that some rows are seen more often than not : for example if you do it with 150 rows: you see 0-99 and then 50-149 and you've seen 2 times more often rows 50-99\n\n## Best solution I found\n\nUse overlapping sliding windows while computing weights such that **each row has exactly a weight one over an entire epoch**. The maths are a bit complex so I'll give the formula :\n\n```python\ndef get_list_of_positions_of_one_sequence(length: int, max_length: int, overlap: int):\n    \"\"\"Return the start position and the associated weight\"\"\"\n    from math import ceil, floor\n    weights = list()\n    for i in range(length):\n        weight = 1 + min(floor(i/overlap), max(0, ceil((length-max_length)/overlap))) -\\\n                 max(0, ceil((i - (max_length - 1)) / overlap))\n        weights.append(weight)\n    weights = 1/np.array(weights)\n    start = 0\n    while True:\n        yield start, weights[start:start+max_length]\n        if start+max_length>=length:\n            break\n        start += overlap\n\n```\n\nIf you have length interactions and using max_length = 100 and an overlap of 50 (you translate of 50 interactions every time), the get_list_of_positions_of_one_sequence(length, 100, 50) will yield the start position and the associated 100 weights you can use.",
    "1119256": "Sampling strategy:\n\n- Take all user ids in training data.\n- Compute their lengths and cap the result to 500.\n- Scale the prior to make them sum 1 (probability distribution).\n- Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.\n- Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.\n- Repeat each epoch.\n\nWhy?\n\n- If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times.\n- Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.\n\n(Extensive discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632))",
    "1119255": "Thank you for sharing, that's really nice. I want to point out that, @claverru  and me also have some similar approach to this issue [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1116239) - despite some differences.",
    "1119250": "I'm at the improvement I level at the moment. \n\nBest solution sound interesting, not sure how much of a gain would that bring. Can you share score of before and after for motivation :D",
    "1127340": "Brilliant. Deserves more upvotes. One quick question though. What is wrong with the strategy of breaking a user>100 timesteps into chunks of 100 as most of the Sota papers have done (solution 1 by itself). Would this perform worse than the improvement suggested?\n\nof course claveru point is also is much valid. There are many starter sequences (student dropouts) and we would need to find out simple ways to overcome that.",
    "1120689": "Just curious, wouldn't it be possible for the model to learn those common sequences as well wich overlap b/w the sub-sequences from the same user?",
    "1119967": "@rodolphelampe - why your LB rank is not shown? You have no submission so far?",
    "1119279": ""
  }
}