{
  "id": 206279,
  "title": "Helpful Observations",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206279",
  "author_name": "Bluefool",
  "post_date": "2020-12-23T21:15:38.874000",
  "votes": 84,
  "comment_count": 35,
  "views": 0,
  "content": "<p>I wanted to enter this competition and have worked on it for a few weeks. However, my software engineering isn't good enough and I will not be able to process the rows in sufficient time. I'm therefore giving up and haven't made a submission.</p>\n<p>Here are some titbits:</p>\n<ol>\n<li><p>My features :<br>\npd.DataFrame(columns=['index', 'row_id', 'timestamp', 'user_id', 'answered_correctly',<br>\n   'number_lectures_attended', 'question_id', 'bundle_id', 'part', 'tags',<br>\n   'no_of_tags_in_question', 'no_questions_in_bundle',<br>\n   'no_answers_in_question', 'no_of_tags_with_lectures',<br>\n   'percent_of_tags_with_lectures', 'cluster', 'average_bundle_time',<br>\n   'average_question_time', 'question_id_mean', 'bundle_id_mean',<br>\n   'part_mean', 'tags_mean', 'cluster_mean', 'average_bundle_time_mean',<br>\n   'average_question_time_mean', 'question_rank', <br>\n 'user_count', 'user_mean', 'user_correct', 'user_part_count', <br>\n 'user_part_mean', 'user_part_correct', 'user_cluster_count', 'user_cluster_mean', 'user_cluster_correct',<br>\n    'first_bundle',  'first_bundle_user_count', 'first_cluster', 'first_cluster_user_count',<br>\n 'last_timestamp', 'last_timestamp_incorrect'])</p></li>\n<li><p>I preprocessed the questions and got mean scores, number of answers in question (using training data) etc. I then bucketed them. When there's 100 million rows in training and 2 million in test, I doubt the average for bundles, time, questions would change by updating them. I dealt with them as if knowing a question was hard or not as a priori. I just had this as a csv to be used by the kernel. These values are not updated by the model - they are fixed.</p></li>\n<li><p>I used the hashing trick for ftrl for user_id, question_id, bundle id etc</p></li>\n<li><p>I binarised the tags and clustered them into 20 clusters. I also ftrl hashed the tags</p></li>\n<li><p>For training, update after a user bundle not a user question. A lot of the scripts update after each question but in the test data , you will only have the correct answer after the whole bundle (so can't use pandas shift etc)</p></li>\n<li><p>Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.</p></li>\n</ol>\n<p>Number 6 is important and should boost your score</p>\n<p>Merry Christmas<br>\nBluefool</p>",
  "messages": [
    {
      "id": 1124357,
      "postDate": "2020-12-23T21:15:38.873Z",
      "content": "<p>I wanted to enter this competition and have worked on it for a few weeks. However, my software engineering isn't good enough and I will not be able to process the rows in sufficient time. I'm therefore giving up and haven't made a submission.</p>\n<p>Here are some titbits:</p>\n<ol>\n<li><p>My features :<br>\npd.DataFrame(columns=['index', 'row_id', 'timestamp', 'user_id', 'answered_correctly',<br>\n   'number_lectures_attended', 'question_id', 'bundle_id', 'part', 'tags',<br>\n   'no_of_tags_in_question', 'no_questions_in_bundle',<br>\n   'no_answers_in_question', 'no_of_tags_with_lectures',<br>\n   'percent_of_tags_with_lectures', 'cluster', 'average_bundle_time',<br>\n   'average_question_time', 'question_id_mean', 'bundle_id_mean',<br>\n   'part_mean', 'tags_mean', 'cluster_mean', 'average_bundle_time_mean',<br>\n   'average_question_time_mean', 'question_rank', <br>\n 'user_count', 'user_mean', 'user_correct', 'user_part_count', <br>\n 'user_part_mean', 'user_part_correct', 'user_cluster_count', 'user_cluster_mean', 'user_cluster_correct',<br>\n    'first_bundle',  'first_bundle_user_count', 'first_cluster', 'first_cluster_user_count',<br>\n 'last_timestamp', 'last_timestamp_incorrect'])</p></li>\n<li><p>I preprocessed the questions and got mean scores, number of answers in question (using training data) etc. I then bucketed them. When there's 100 million rows in training and 2 million in test, I doubt the average for bundles, time, questions would change by updating them. I dealt with them as if knowing a question was hard or not as a priori. I just had this as a csv to be used by the kernel. These values are not updated by the model - they are fixed.</p></li>\n<li><p>I used the hashing trick for ftrl for user_id, question_id, bundle id etc</p></li>\n<li><p>I binarised the tags and clustered them into 20 clusters. I also ftrl hashed the tags</p></li>\n<li><p>For training, update after a user bundle not a user question. A lot of the scripts update after each question but in the test data , you will only have the correct answer after the whole bundle (so can't use pandas shift etc)</p></li>\n<li><p>Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.</p></li>\n</ol>\n<p>Number 6 is important and should boost your score</p>\n<p>Merry Christmas<br>\nBluefool</p>",
      "rawMarkdown": "I wanted to enter this competition and have worked on it for a few weeks. However, my software engineering isn't good enough and I will not be able to process the rows in sufficient time. I'm therefore giving up and haven't made a submission.\n\nHere are some titbits:\n\n1.  My features :\npd.DataFrame(columns=['index', 'row_id', 'timestamp', 'user_id', 'answered_correctly',\n       'number_lectures_attended', 'question_id', 'bundle_id', 'part', 'tags',\n       'no_of_tags_in_question', 'no_questions_in_bundle',\n       'no_answers_in_question', 'no_of_tags_with_lectures',\n       'percent_of_tags_with_lectures', 'cluster', 'average_bundle_time',\n       'average_question_time', 'question_id_mean', 'bundle_id_mean',\n       'part_mean', 'tags_mean', 'cluster_mean', 'average_bundle_time_mean',\n       'average_question_time_mean', 'question_rank', \n     'user_count', 'user_mean', 'user_correct', 'user_part_count', \n     'user_part_mean', 'user_part_correct', 'user_cluster_count', 'user_cluster_mean', 'user_cluster_correct',\n        'first_bundle',  'first_bundle_user_count', 'first_cluster', 'first_cluster_user_count',\n     'last_timestamp', 'last_timestamp_incorrect'])\n\n2. I preprocessed the questions and got mean scores, number of answers in question (using training data) etc. I then bucketed them. When there's 100 million rows in training and 2 million in test, I doubt the average for bundles, time, questions would change by updating them. I dealt with them as if knowing a question was hard or not as a priori. I just had this as a csv to be used by the kernel. These values are not updated by the model - they are fixed.\n\n3. I used the hashing trick for ftrl for user_id, question_id, bundle id etc\n\n4. I binarised the tags and clustered them into 20 clusters. I also ftrl hashed the tags\n\n5. For training, update after a user bundle not a user question. A lot of the scripts update after each question but in the test data , you will only have the correct answer after the whole bundle (so can't use pandas shift etc)\n\n6. Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.\n\nNumber 6 is important and should boost your score\n\nMerry Christmas\nBluefool\n\n\n\n",
      "votes": 84
    },
    {
      "id": 1125033,
      "postDate": "2020-12-24T10:59:42.093Z",
      "content": "<blockquote>\n  <p>Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.</p>\n</blockquote>\n<p>I am very suspicious about this. 50% of the students in train set have 7900 as the first bundle_id. While for our purposes with this dataset they have anonymized the absolute datetime from us, on the Kaggle / Riiid AIEd side, I am almost all but certain that they will be evaluating on a time-based split. If that is the case, it's possible that early on in the program they had a smaller subset of questions, or less 'AI' implemented, resulting in everyone having the same initial question/bundle. </p>\n<p>While it's possible that someone on the Kaggle or host side thought, \"hey, now that we've masked seasonal effects they can focus just on per user time series\" and then subsequently also prepared the public / private lbs as a random split… it'd strike me as very odd. Even in the saint/saint+ papers, eval was always the most recent data.</p>\n<p>I guess best way to test is to prepare a sub against public lb, see how that handles compared to cv boost if any, and then pray that translates into private lb.</p>",
      "rawMarkdown": "> Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.\n\nI am very suspicious about this. 50% of the students in train set have 7900 as the first bundle_id. While for our purposes with this dataset they have anonymized the absolute datetime from us, on the Kaggle / Riiid AIEd side, I am almost all but certain that they will be evaluating on a time-based split. If that is the case, it's possible that early on in the program they had a smaller subset of questions, or less 'AI' implemented, resulting in everyone having the same initial question/bundle. \n\nWhile it's possible that someone on the Kaggle or host side thought, \"hey, now that we've masked seasonal effects they can focus just on per user time series\" and then subsequently also prepared the public / private lbs as a random split... it'd strike me as very odd. Even in the saint/saint+ papers, eval was always the most recent data.\n\nI guess best way to test is to prepare a sub against public lb, see how that handles compared to cv boost if any, and then pray that translates into private lb.",
      "votes": 4,
      "replies": [
        {
          "id": 1125757,
          "postDate": "2020-12-25T03:17:31.403Z",
          "content": "<p>Has Point#6 above helped anybody? I tried a single first_bundle==7900 feature, a generic first_bundle embedding, and first_bundle embedding where bundles &lt; 40 users all get bucketed into one bin. None of these schemes had a positive impact on my local validation.</p>\n<p>Anyone?</p>\n<p>I haven't attempted first tag clustering, though I have made use of tag clustering features per the public kernels.</p>",
          "rawMarkdown": "Has Point#6 above helped anybody? I tried a single first_bundle==7900 feature, a generic first_bundle embedding, and first_bundle embedding where bundles < 40 users all get bucketed into one bin. None of these schemes had a positive impact on my local validation.\n\nAnyone?\n\nI haven't attempted first tag clustering, though I have made use of tag clustering features per the public kernels."
        },
        {
          "id": 1125924,
          "postDate": "2020-12-25T07:19:43.373Z",
          "content": "<p>Did tag clusters help you?</p>",
          "rawMarkdown": "Did tag clusters help you?"
        },
        {
          "id": 1126261,
          "postDate": "2020-12-25T13:22:16.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/nicohrubec\" target=\"_blank\">@nicohrubec</a> please see my comment on <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a>'s kernel <a href=\"https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags/comments#1124399\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "@nicohrubec please see my comment on @spacelx's kernel [here](https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags/comments#1124399).",
          "votes": 1
        },
        {
          "id": 1126269,
          "postDate": "2020-12-25T13:29:02.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Interesting. Thank you!</p>",
          "rawMarkdown": "@authman Interesting. Thank you!"
        }
      ]
    },
    {
      "id": 1125462,
      "postDate": "2020-12-24T18:18:44.810Z",
      "content": "<p>Currently, I use 69 features, and the submission running time is &lt; 40 mins which means my codes are deeply optimized with many tricks I have never seen in the public notebooks. Thanks for your ideas, I will try your ideas this Sunday after finishing one certification exam. If it boosts my score, I will team up with you.</p>",
      "rawMarkdown": "Currently, I use 69 features, and the submission running time is < 40 mins which means my codes are deeply optimized with many tricks I have never seen in the public notebooks. Thanks for your ideas, I will try your ideas this Sunday after finishing one certification exam. If it boosts my score, I will team up with you.",
      "votes": 2
    },
    {
      "id": 1131847,
      "postDate": "2020-12-30T03:07:21.883Z",
      "content": "<p>Thanks for the tips! I am just wondering what does hashing trick means?</p>",
      "rawMarkdown": "Thanks for the tips! I am just wondering what does hashing trick means?",
      "votes": 1,
      "replies": [
        {
          "id": 1133211,
          "postDate": "2020-12-31T02:57:24.613Z",
          "content": "<p>I wonder about that too</p>",
          "rawMarkdown": "I wonder about that too\n"
        },
        {
          "id": 1133572,
          "postDate": "2020-12-31T10:32:36.367Z",
          "content": "<p>Here is the sklearn implementation and description:<br>\n<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.FeatureHasher.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.FeatureHasher.html</a></p>",
          "rawMarkdown": "Here is the sklearn implementation and description:\n[https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.FeatureHasher.html](https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.FeatureHasher.html)",
          "votes": 4
        },
        {
          "id": 1133595,
          "postDate": "2020-12-31T11:13:14.173Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> for the link. One follow up question: How can we use hashing here? (For ftrl) </p>",
          "rawMarkdown": "Thanks @bowaka for the link. One follow up question: How can we use hashing here? (For ftrl) "
        }
      ]
    },
    {
      "id": 1128999,
      "postDate": "2020-12-28T01:30:31.850Z",
      "content": "<p>I added <code>starter bundle</code> the AUC on the training set increased from 0.810418 to 0.818764 but the AUC on the validation set decreased from 0.787183 to 0.783996</p>",
      "rawMarkdown": "I added `starter bundle` the AUC on the training set increased from 0.810418 to 0.818764 but the AUC on the validation set decreased from 0.787183 to 0.783996",
      "votes": 1,
      "replies": [
        {
          "id": 1129670,
          "postDate": "2020-12-28T14:00:46.090Z",
          "content": "<p>Thanks for reporting and corroborating my assumption. If 1/2 the questions share the same starter bundle (and we already know that SANTA app starts with a series of diagnostic questions) then it doesn't make sense to assume there is signal in the first bundle a user is exposed to. It'd be a different story if there was a pre-pre-test that then bucketed people and had them start with questions based on accessed skill level. But we aren't told when the diagnostic questions end and when the real questions begin to make that calculation.</p>",
          "rawMarkdown": "Thanks for reporting and corroborating my assumption. If 1/2 the questions share the same starter bundle (and we already know that SANTA app starts with a series of diagnostic questions) then it doesn't make sense to assume there is signal in the first bundle a user is exposed to. It'd be a different story if there was a pre-pre-test that then bucketed people and had them start with questions based on accessed skill level. But we aren't told when the diagnostic questions end and when the real questions begin to make that calculation."
        },
        {
          "id": 1129763,
          "postDate": "2020-12-28T14:32:18.790Z",
          "content": "<p>The diagnostic questions dont have explanation..So maybe we keep rolling till we get questions with explanation and then we know that diagnostic questions have ended. But this is tricky biz for sure..</p>\n<p>But Isn't 5 a brilliant insight? wonder why no-one is talking about it. </p>",
          "rawMarkdown": "The diagnostic questions dont have explanation..So maybe we keep rolling till we get questions with explanation and then we know that diagnostic questions have ended. But this is tricky biz for sure..\n\nBut Isn't 5 a brilliant insight? wonder why no-one is talking about it. "
        },
        {
          "id": 1129771,
          "postDate": "2020-12-28T14:39:32.137Z",
          "content": "<p>It is but it's a pain to implement. Maybe for top lb scorers it'll help curb overfitting a bit. But if anything, it's a regularizer not a feature to boost score. <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> mentioned doing something similar by means of using the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620\" target=\"_blank\">attention mask itself</a>, e.g. doing the update per-question, but only allowing the transformer to attend to out-of-bundle questions. I thought that was a more elegant implementation to accomplish more or less the same purpose. Coincidentally, his post also didn't get much 'attention' :P.</p>",
          "rawMarkdown": "It is but it's a pain to implement. Maybe for top lb scorers it'll help curb overfitting a bit. But if anything, it's a regularizer not a feature to boost score. @rodolphelampe mentioned doing something similar by means of using the [attention mask itself](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620), e.g. doing the update per-question, but only allowing the transformer to attend to out-of-bundle questions. I thought that was a more elegant implementation to accomplish more or less the same purpose. Coincidentally, his post also didn't get much 'attention' :P."
        },
        {
          "id": 1130146,
          "postDate": "2020-12-28T19:05:27.150Z",
          "content": "<p>Not sure if its the best approach, but for item 5, after all your feature extraction and updates, you can do something like:</p>\n<p><code>train.groupby([\"user_id\", \"task_container_id\"])[\"your_feature\"].transform(\"first\")</code></p>",
          "rawMarkdown": "Not sure if its the best approach, but for item 5, after all your feature extraction and updates, you can do something like:\n\n`train.groupby([\"user_id\", \"task_container_id\"])[\"your_feature\"].transform(\"first\")`",
          "votes": 3
        },
        {
          "id": 1130160,
          "postDate": "2020-12-28T19:19:13.803Z",
          "content": "<p>I personnaly create a temporary list, in which I put the rows from the same bundle_id. I then loop over that list to create the features, and once all features are created, I reloop on the list to update the cache files. It allows me as well to consider bundle_size as a feature.</p>",
          "rawMarkdown": "I personnaly create a temporary list, in which I put the rows from the same bundle_id. I then loop over that list to create the features, and once all features are created, I reloop on the list to update the cache files. It allows me as well to consider bundle_size as a feature.",
          "votes": 1
        },
        {
          "id": 1130247,
          "postDate": "2020-12-28T21:47:52.487Z",
          "content": "<p>I take my comment back. I highly recommend everyone try out <a href=\"https://www.kaggle.com/jsaguiar\" target=\"_blank\">@jsaguiar</a>'s great tip above:</p>\n<blockquote>\n  <p><code>train.groupby([\"user_id\", \"task_container_id\"])[\"your_feature\"].transform(\"first\")</code></p>\n</blockquote>",
          "rawMarkdown": "I take my comment back. I highly recommend everyone try out @jsaguiar's great tip above:\n\n> `train.groupby([\"user_id\", \"task_container_id\"])[\"your_feature\"].transform(\"first\")`\n",
          "votes": 1
        },
        {
          "id": 1130430,
          "postDate": "2020-12-29T03:34:29.247Z",
          "content": "<p>elegant!<br>\nalso intuitively this should help for seq models surely</p>",
          "rawMarkdown": "elegant!\nalso intuitively this should help for seq models surely"
        },
        {
          "id": 1130642,
          "postDate": "2020-12-29T07:41:17.470Z",
          "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> how you calculate the cv? all sub item(not include padding value) in sequence AUC? or just last sub item in each sequence AUC?</p>",
          "rawMarkdown": "@wuwenmin how you calculate the cv? all sub item(not include padding value) in sequence AUC? or just last sub item in each sequence AUC?\n"
        },
        {
          "id": 1130685,
          "postDate": "2020-12-29T08:30:02.737Z",
          "content": "<p>I use timestamp check if the question in the same bundle,update dict while different bundle.it's useful for my LGBM</p>",
          "rawMarkdown": "I use timestamp check if the question in the same bundle,update dict while different bundle.it's useful for my LGBM"
        }
      ]
    },
    {
      "id": 1124719,
      "postDate": "2020-12-24T06:54:53.563Z",
      "content": "<p>Hi Bluefool, </p>\n<p>Thanks for sharing. I have one question on this part: </p>\n<ol>\n<li>Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and <strong>extract all the users with the same start</strong>.</li>\n</ol>\n<p>What do you mean when you say extract all the users with the same start? </p>\n<p>Merry Christmas, <br>\nCakey</p>",
      "rawMarkdown": "Hi Bluefool, \n\nThanks for sharing. I have one question on this part: \n\n6. Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and **extract all the users with the same start**.\n\nWhat do you mean when you say extract all the users with the same start? \n\nMerry Christmas, \nCakey\n",
      "votes": 1,
      "replies": [
        {
          "id": 1124929,
          "postDate": "2020-12-24T09:24:09.607Z",
          "content": "<p>possibly he means that is one way of clustering students?</p>",
          "rawMarkdown": "possibly he means that is one way of clustering students?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1124377,
      "postDate": "2020-12-23T21:46:00.943Z",
      "content": "<p>Thank you, that's very kind to share so much !<br>\nI feel bad for you as you must have spend a lot of time on this competition. I spent a lot of time improving the computation time and I think we could help you. Do you have any script to share that shows where it is slow ?<br>\nI used viztracer which is an awesome library for profiling your code (using the chrome tracing tool), you could share that or I don't know what but feel free to share your difficulties and there might be solutions.</p>",
      "rawMarkdown": "Thank you, that's very kind to share so much !\nI feel bad for you as you must have spend a lot of time on this competition. I spent a lot of time improving the computation time and I think we could help you. Do you have any script to share that shows where it is slow ?\nI used viztracer which is an awesome library for profiling your code (using the chrome tracing tool), you could share that or I don't know what but feel free to share your difficulties and there might be solutions.",
      "votes": 1
    },
    {
      "id": 1124645,
      "postDate": "2020-12-24T05:41:04.127Z",
      "content": "<p>Thank you for your sharing,<br>\nCould you explain what does it mean</p>\n<ol>\n<li>'same amount of questions answered, same scores.'</li>\n</ol>\n<p>I've found that users whose first bundle_id is same have same subsequent content_ids (thanks to you)<br>\nsame amount of questions answered, same scores mean this?</p>\n<p>Wish you Merry Christmas<br>\nBlueBird</p>",
      "rawMarkdown": "Thank you for your sharing,\nCould you explain what does it mean\n6. 'same amount of questions answered, same scores.'\n\nI've found that users whose first bundle_id is same have same subsequent content_ids (thanks to you)\nsame amount of questions answered, same scores mean this?\n\nWish you Merry Christmas\nBlueBird",
      "votes": 2
    },
    {
      "id": 1130989,
      "postDate": "2020-12-29T13:32:00.203Z",
      "content": "<p>Anybody running into this error: <a href=\"https://github.com/pytorch/pytorch/pull/24888\" target=\"_blank\">https://github.com/pytorch/pytorch/pull/24888</a> <br>\nThis is during applying padding mask. Any suggestions or reference on padding masks welcome<br>\nNote: This is not peek-ahead mask. That is well implemented in the public kernels</p>",
      "rawMarkdown": "Anybody running into this error: https://github.com/pytorch/pytorch/pull/24888 \nThis is during applying padding mask. Any suggestions or reference on padding masks welcome\nNote: This is not peek-ahead mask. That is well implemented in the public kernels"
    },
    {
      "id": 1127018,
      "postDate": "2020-12-26T07:10:40.707Z",
      "content": "<p>Thanks for sharing and merry Christmas <a href=\"https://www.kaggle.com/domcastro\" target=\"_blank\">@domcastro</a> </p>",
      "rawMarkdown": "Thanks for sharing and merry Christmas @domcastro "
    },
    {
      "id": 1125578,
      "postDate": "2020-12-24T20:36:03.473Z",
      "content": "<p>sir, are there any good links for ftrl trick and I want to learn about it.</p>",
      "rawMarkdown": "sir, are there any good links for ftrl trick and I want to learn about it.",
      "replies": [
        {
          "id": 1125600,
          "postDate": "2020-12-24T21:07:12.947Z",
          "content": "<p><a href=\"https://research.google.com/pubs/archive/41159.pdf\" target=\"_blank\">https://research.google.com/pubs/archive/41159.pdf</a></p>\n<p>It is implemented as an optimizer in Keras.</p>",
          "rawMarkdown": "https://research.google.com/pubs/archive/41159.pdf\n\nIt is implemented as an optimizer in Keras.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1125406,
      "postDate": "2020-12-24T17:30:48.073Z",
      "content": "<p>Also, how many data fields per user do you need to maintain for your model? If it is a reasonable number (&lt;100 or so), then it should not be difficult to make it run in 9 hours. But if it is a large number (like 10,000 or more), then it may not be possible to make it work at all.</p>",
      "rawMarkdown": "Also, how many data fields per user do you need to maintain for your model? If it is a reasonable number (<100 or so), then it should not be difficult to make it run in 9 hours. But if it is a large number (like 10,000 or more), then it may not be possible to make it work at all."
    },
    {
      "id": 1125399,
      "postDate": "2020-12-24T17:20:32.863Z",
      "content": "<p>What is your CV? If it is high enough (&gt;0.800) you can team up with somebody who can help you optimize the submission code.</p>",
      "rawMarkdown": "What is your CV? If it is high enough (>0.800) you can team up with somebody who can help you optimize the submission code.",
      "replies": [
        {
          "id": 1130569,
          "postDate": "2020-12-29T06:04:22.007Z",
          "content": "<p><a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> would u like to team up. our standing is just contributed by Saint model of 78.6 &amp; basic lgb 77.2, we have built draft saint plus with 78.8 lb . we are making it better like saint. We target for gold  let me know if you want to team up ..With our pipeline scores can be boosted further. we will have  an lgb model  with score of 79.2 . Towards the end we can get more boost and we will lend up in gold zone.</p>",
          "rawMarkdown": "@ymatioun would u like to team up. our standing is just contributed by Saint model of 78.6 & basic lgb 77.2, we have built draft saint plus with 78.8 lb . we are making it better like saint. We target for gold  let me know if you want to team up ..With our pipeline scores can be boosted further. we will have  an lgb model  with score of 79.2 . Towards the end we can get more boost and we will lend up in gold zone.",
          "votes": -1
        },
        {
          "id": 1130774,
          "postDate": "2020-12-29T09:52:20.720Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1124508,
      "postDate": "2020-12-24T02:14:13.123Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1128988,
      "postDate": "2020-12-28T00:53:00.843Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1126057,
      "postDate": "2020-12-25T09:55:12.647Z",
      "content": "<p>Thank you for your selfless sharing！</p>",
      "rawMarkdown": "Thank you for your selfless sharing！",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1125033,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-12-24T10:59:42.093000",
      "content": "<blockquote>\n  <p>Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.</p>\n</blockquote>\n<p>I am very suspicious about this. 50% of the students in train set have 7900 as the first bundle_id. While for our purposes with this dataset they have anonymized the absolute datetime from us, on the Kaggle / Riiid AIEd side, I am almost all but certain that they will be evaluating on a time-based split. If that is the case, it's possible that early on in the program they had a smaller subset of questions, or less 'AI' implemented, resulting in everyone having the same initial question/bundle. </p>\n<p>While it's possible that someone on the Kaggle or host side thought, \"hey, now that we've masked seasonal effects they can focus just on per user time series\" and then subsequently also prepared the public / private lbs as a random split… it'd strike me as very odd. Even in the saint/saint+ papers, eval was always the most recent data.</p>\n<p>I guess best way to test is to prepare a sub against public lb, see how that handles compared to cv boost if any, and then pray that translates into private lb.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1125757,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-25T03:17:31.403000",
          "content": "<p>Has Point#6 above helped anybody? I tried a single first_bundle==7900 feature, a generic first_bundle embedding, and first_bundle embedding where bundles &lt; 40 users all get bucketed into one bin. None of these schemes had a positive impact on my local validation.</p>\n<p>Anyone?</p>\n<p>I haven't attempted first tag clustering, though I have made use of tag clustering features per the public kernels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1125924,
          "author_name": "hrunic",
          "author_url": "",
          "post_date": "2020-12-25T07:19:43.373000",
          "content": "<p>Did tag clusters help you?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1126261,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-25T13:22:16.690000",
          "content": "<p><a href=\"https://www.kaggle.com/nicohrubec\" target=\"_blank\">@nicohrubec</a> please see my comment on <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a>'s kernel <a href=\"https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags/comments#1124399\" target=\"_blank\">here</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1126269,
          "author_name": "hrunic",
          "author_url": "",
          "post_date": "2020-12-25T13:29:02.100000",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Interesting. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1125462,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2020-12-24T18:18:44.810000",
      "content": "<p>Currently, I use 69 features, and the submission running time is &lt; 40 mins which means my codes are deeply optimized with many tricks I have never seen in the public notebooks. Thanks for your ideas, I will try your ideas this Sunday after finishing one certification exam. If it boosts my score, I will team up with you.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1131847,
      "author_name": "SHYJohn",
      "author_url": "",
      "post_date": "2020-12-30T03:07:21.883000",
      "content": "<p>Thanks for the tips! I am just wondering what does hashing trick means?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1133211,
          "author_name": "jwc",
          "author_url": "",
          "post_date": "2020-12-31T02:57:24.613000",
          "content": "<p>I wonder about that too</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133572,
          "author_name": "Jacky",
          "author_url": "",
          "post_date": "2020-12-31T10:32:36.367000",
          "content": "<p>Here is the sklearn implementation and description:<br>\n<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.FeatureHasher.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.FeatureHasher.html</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1133595,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-31T11:13:14.173000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> for the link. One follow up question: How can we use hashing here? (For ftrl) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1128999,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2020-12-28T01:30:31.850000",
      "content": "<p>I added <code>starter bundle</code> the AUC on the training set increased from 0.810418 to 0.818764 but the AUC on the validation set decreased from 0.787183 to 0.783996</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1129670,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-28T14:00:46.090000",
          "content": "<p>Thanks for reporting and corroborating my assumption. If 1/2 the questions share the same starter bundle (and we already know that SANTA app starts with a series of diagnostic questions) then it doesn't make sense to assume there is signal in the first bundle a user is exposed to. It'd be a different story if there was a pre-pre-test that then bucketed people and had them start with questions based on accessed skill level. But we aren't told when the diagnostic questions end and when the real questions begin to make that calculation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129763,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-28T14:32:18.790000",
          "content": "<p>The diagnostic questions dont have explanation..So maybe we keep rolling till we get questions with explanation and then we know that diagnostic questions have ended. But this is tricky biz for sure..</p>\n<p>But Isn't 5 a brilliant insight? wonder why no-one is talking about it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129771,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-28T14:39:32.137000",
          "content": "<p>It is but it's a pain to implement. Maybe for top lb scorers it'll help curb overfitting a bit. But if anything, it's a regularizer not a feature to boost score. <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> mentioned doing something similar by means of using the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620\" target=\"_blank\">attention mask itself</a>, e.g. doing the update per-question, but only allowing the transformer to attend to out-of-bundle questions. I thought that was a more elegant implementation to accomplish more or less the same purpose. Coincidentally, his post also didn't get much 'attention' :P.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130146,
          "author_name": "Aguiar",
          "author_url": "",
          "post_date": "2020-12-28T19:05:27.150000",
          "content": "<p>Not sure if its the best approach, but for item 5, after all your feature extraction and updates, you can do something like:</p>\n<p><code>train.groupby([\"user_id\", \"task_container_id\"])[\"your_feature\"].transform(\"first\")</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1130160,
          "author_name": "Jacky",
          "author_url": "",
          "post_date": "2020-12-28T19:19:13.803000",
          "content": "<p>I personnaly create a temporary list, in which I put the rows from the same bundle_id. I then loop over that list to create the features, and once all features are created, I reloop on the list to update the cache files. It allows me as well to consider bundle_size as a feature.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1130247,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-28T21:47:52.487000",
          "content": "<p>I take my comment back. I highly recommend everyone try out <a href=\"https://www.kaggle.com/jsaguiar\" target=\"_blank\">@jsaguiar</a>'s great tip above:</p>\n<blockquote>\n  <p><code>train.groupby([\"user_id\", \"task_container_id\"])[\"your_feature\"].transform(\"first\")</code></p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1130430,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-29T03:34:29.247000",
          "content": "<p>elegant!<br>\nalso intuitively this should help for seq models surely</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130642,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-12-29T07:41:17.470000",
          "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> how you calculate the cv? all sub item(not include padding value) in sequence AUC? or just last sub item in each sequence AUC?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130685,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-29T08:30:02.737000",
          "content": "<p>I use timestamp check if the question in the same bundle,update dict while different bundle.it's useful for my LGBM</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1124719,
      "author_name": "Cakey",
      "author_url": "",
      "post_date": "2020-12-24T06:54:53.563000",
      "content": "<p>Hi Bluefool, </p>\n<p>Thanks for sharing. I have one question on this part: </p>\n<ol>\n<li>Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and <strong>extract all the users with the same start</strong>.</li>\n</ol>\n<p>What do you mean when you say extract all the users with the same start? </p>\n<p>Merry Christmas, <br>\nCakey</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1124929,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-24T09:24:09.607000",
          "content": "<p>possibly he means that is one way of clustering students?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1124377,
      "author_name": "Rodolphe Lampe",
      "author_url": "",
      "post_date": "2020-12-23T21:46:00.943000",
      "content": "<p>Thank you, that's very kind to share so much !<br>\nI feel bad for you as you must have spend a lot of time on this competition. I spent a lot of time improving the computation time and I think we could help you. Do you have any script to share that shows where it is slow ?<br>\nI used viztracer which is an awesome library for profiling your code (using the chrome tracing tool), you could share that or I don't know what but feel free to share your difficulties and there might be solutions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1124645,
      "author_name": "Potential",
      "author_url": "",
      "post_date": "2020-12-24T05:41:04.127000",
      "content": "<p>Thank you for your sharing,<br>\nCould you explain what does it mean</p>\n<ol>\n<li>'same amount of questions answered, same scores.'</li>\n</ol>\n<p>I've found that users whose first bundle_id is same have same subsequent content_ids (thanks to you)<br>\nsame amount of questions answered, same scores mean this?</p>\n<p>Wish you Merry Christmas<br>\nBlueBird</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1130989,
      "author_name": "Allohvk",
      "author_url": "",
      "post_date": "2020-12-29T13:32:00.203000",
      "content": "<p>Anybody running into this error: <a href=\"https://github.com/pytorch/pytorch/pull/24888\" target=\"_blank\">https://github.com/pytorch/pytorch/pull/24888</a> <br>\nThis is during applying padding mask. Any suggestions or reference on padding masks welcome<br>\nNote: This is not peek-ahead mask. That is well implemented in the public kernels</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1127018,
      "author_name": "Saurabh Shahane",
      "author_url": "",
      "post_date": "2020-12-26T07:10:40.707000",
      "content": "<p>Thanks for sharing and merry Christmas <a href=\"https://www.kaggle.com/domcastro\" target=\"_blank\">@domcastro</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1125578,
      "author_name": "Zhenghan Chen",
      "author_url": "",
      "post_date": "2020-12-24T20:36:03.473000",
      "content": "<p>sir, are there any good links for ftrl trick and I want to learn about it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1125600,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-24T21:07:12.947000",
          "content": "<p><a href=\"https://research.google.com/pubs/archive/41159.pdf\" target=\"_blank\">https://research.google.com/pubs/archive/41159.pdf</a></p>\n<p>It is implemented as an optimizer in Keras.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1125406,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2020-12-24T17:30:48.073000",
      "content": "<p>Also, how many data fields per user do you need to maintain for your model? If it is a reasonable number (&lt;100 or so), then it should not be difficult to make it run in 9 hours. But if it is a large number (like 10,000 or more), then it may not be possible to make it work at all.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1125399,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2020-12-24T17:20:32.863000",
      "content": "<p>What is your CV? If it is high enough (&gt;0.800) you can team up with somebody who can help you optimize the submission code.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1130569,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-29T06:04:22.007000",
          "content": "<p><a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> would u like to team up. our standing is just contributed by Saint model of 78.6 &amp; basic lgb 77.2, we have built draft saint plus with 78.8 lb . we are making it better like saint. We target for gold  let me know if you want to team up ..With our pipeline scores can be boosted further. we will have  an lgb model  with score of 79.2 . Towards the end we can get more boost and we will lend up in gold zone.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1130774,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-29T09:52:20.720000",
          "content": "",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 1124508,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-24T02:14:13.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1128988,
      "author_name": "Lokesh",
      "author_url": "",
      "post_date": "2020-12-28T00:53:00.843000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1126057,
      "author_name": "Yi Song",
      "author_url": "",
      "post_date": "2020-12-25T09:55:12.647000",
      "content": "<p>Thank you for your selfless sharing！</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1124357": "I wanted to enter this competition and have worked on it for a few weeks. However, my software engineering isn't good enough and I will not be able to process the rows in sufficient time. I'm therefore giving up and haven't made a submission.\n\nHere are some titbits:\n\n1.  My features :\npd.DataFrame(columns=['index', 'row_id', 'timestamp', 'user_id', 'answered_correctly',\n       'number_lectures_attended', 'question_id', 'bundle_id', 'part', 'tags',\n       'no_of_tags_in_question', 'no_questions_in_bundle',\n       'no_answers_in_question', 'no_of_tags_with_lectures',\n       'percent_of_tags_with_lectures', 'cluster', 'average_bundle_time',\n       'average_question_time', 'question_id_mean', 'bundle_id_mean',\n       'part_mean', 'tags_mean', 'cluster_mean', 'average_bundle_time_mean',\n       'average_question_time_mean', 'question_rank', \n     'user_count', 'user_mean', 'user_correct', 'user_part_count', \n     'user_part_mean', 'user_part_correct', 'user_cluster_count', 'user_cluster_mean', 'user_cluster_correct',\n        'first_bundle',  'first_bundle_user_count', 'first_cluster', 'first_cluster_user_count',\n     'last_timestamp', 'last_timestamp_incorrect'])\n\n2. I preprocessed the questions and got mean scores, number of answers in question (using training data) etc. I then bucketed them. When there's 100 million rows in training and 2 million in test, I doubt the average for bundles, time, questions would change by updating them. I dealt with them as if knowing a question was hard or not as a priori. I just had this as a csv to be used by the kernel. These values are not updated by the model - they are fixed.\n\n3. I used the hashing trick for ftrl for user_id, question_id, bundle id etc\n\n4. I binarised the tags and clustered them into 20 clusters. I also ftrl hashed the tags\n\n5. For training, update after a user bundle not a user question. A lot of the scripts update after each question but in the test data , you will only have the correct answer after the whole bundle (so can't use pandas shift etc)\n\n6. Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.\n\nNumber 6 is important and should boost your score\n\nMerry Christmas\nBluefool\n\n\n\n",
    "1125033": "> Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and extract all the users with the same start.\n\nI am very suspicious about this. 50% of the students in train set have 7900 as the first bundle_id. While for our purposes with this dataset they have anonymized the absolute datetime from us, on the Kaggle / Riiid AIEd side, I am almost all but certain that they will be evaluating on a time-based split. If that is the case, it's possible that early on in the program they had a smaller subset of questions, or less 'AI' implemented, resulting in everyone having the same initial question/bundle. \n\nWhile it's possible that someone on the Kaggle or host side thought, \"hey, now that we've masked seasonal effects they can focus just on per user time series\" and then subsequently also prepared the public / private lbs as a random split... it'd strike me as very odd. Even in the saint/saint+ papers, eval was always the most recent data.\n\nI guess best way to test is to prepare a sub against public lb, see how that handles compared to cv boost if any, and then pray that translates into private lb.",
    "1125462": "Currently, I use 69 features, and the submission running time is < 40 mins which means my codes are deeply optimized with many tricks I have never seen in the public notebooks. Thanks for your ideas, I will try your ideas this Sunday after finishing one certification exam. If it boosts my score, I will team up with you.",
    "1131847": "Thanks for the tips! I am just wondering what does hashing trick means?",
    "1128999": "I added `starter bundle` the AUC on the training set increased from 0.810418 to 0.818764 but the AUC on the validation set decreased from 0.787183 to 0.783996",
    "1124719": "Hi Bluefool, \n\nThanks for sharing. I have one question on this part: \n\n6. Look at the starter bundle of each user. 7900 is very common. You will see that the bundle that the user starts has a subsequent pattern - same amount of questions answered, same scores. Therefore I capture the users very first bundle and first cluster (of tags) and **extract all the users with the same start**.\n\nWhat do you mean when you say extract all the users with the same start? \n\nMerry Christmas, \nCakey\n",
    "1124377": "Thank you, that's very kind to share so much !\nI feel bad for you as you must have spend a lot of time on this competition. I spent a lot of time improving the computation time and I think we could help you. Do you have any script to share that shows where it is slow ?\nI used viztracer which is an awesome library for profiling your code (using the chrome tracing tool), you could share that or I don't know what but feel free to share your difficulties and there might be solutions.",
    "1124645": "Thank you for your sharing,\nCould you explain what does it mean\n6. 'same amount of questions answered, same scores.'\n\nI've found that users whose first bundle_id is same have same subsequent content_ids (thanks to you)\nsame amount of questions answered, same scores mean this?\n\nWish you Merry Christmas\nBlueBird",
    "1130989": "Anybody running into this error: https://github.com/pytorch/pytorch/pull/24888 \nThis is during applying padding mask. Any suggestions or reference on padding masks welcome\nNote: This is not peek-ahead mask. That is well implemented in the public kernels",
    "1127018": "Thanks for sharing and merry Christmas @domcastro ",
    "1125578": "sir, are there any good links for ftrl trick and I want to learn about it.",
    "1125406": "Also, how many data fields per user do you need to maintain for your model? If it is a reasonable number (<100 or so), then it should not be difficult to make it run in 9 hours. But if it is a large number (like 10,000 or more), then it may not be possible to make it work at all.",
    "1125399": "What is your CV? If it is high enough (>0.800) you can team up with somebody who can help you optimize the submission code.",
    "1124508": "",
    "1128988": "Thanks for sharing.",
    "1126057": "Thank you for your selfless sharing！"
  }
}