{
  "id": 209396,
  "title": "用4G内存处理200+的特征",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209396",
  "author_name": "",
  "post_date": "2021-01-07T13:20:57.061132700Z",
  "votes": -22,
  "comment_count": 30,
  "views": 0,
  "content": "<p>每个用户使用一个User 对象来动态更新状态特征。使用遍历的方法更新用户对象以及输出特征。<br>\n用一个user_feat_dict 来存储所有用户和特征的关系。这是一个很大的字典。<br>\n这个字典由两部分组成，</p>\n<p>第一部分。使用 sqlitedict 调用持久化在磁盘里的特征字典 （db_user_dict）<br>\n第二部分。根据测试集的用户id 动态加载相应的用户特征到内存里。</p>\n<p>由于测试集中的用户很少。所以只提取测试集的用户id缓存到内存里。<br>\n这样40G的线下特征文件，线上实际上内存占用约4G左右。</p>\n<pre><code>#用户字典\nuser_dict = {\n}\ndef get_user(user_id,user_dict):\n    # 如果缓存了\n    try:\n        return user_dict[user_id]\n    except:\n        #没有缓存，看有没有历史\n        try:\n            user_dict[user_id] = db_user_dict_list[user_id % 10][user_id]\n        except:\n            user_dict[user_id] = User()\n        return user_dict[user_id]\n</code></pre>",
  "messages": [
    {
      "id": "1142560",
      "postDate": "01/07/2021 13:20:57",
      "content": "<p>每个用户使用一个User 对象来动态更新状态特征。使用遍历的方法更新用户对象以及输出特征。<br>\n用一个user_feat_dict 来存储所有用户和特征的关系。这是一个很大的字典。<br>\n这个字典由两部分组成，</p>\n<p>第一部分。使用 sqlitedict 调用持久化在磁盘里的特征字典 （db_user_dict）<br>\n第二部分。根据测试集的用户id 动态加载相应的用户特征到内存里。</p>\n<p>由于测试集中的用户很少。所以只提取测试集的用户id缓存到内存里。<br>\n这样40G的线下特征文件，线上实际上内存占用约4G左右。</p>\n<pre><code>#用户字典\nuser_dict = {\n}\ndef get_user(user_id,user_dict):\n    # 如果缓存了\n    try:\n        return user_dict[user_id]\n    except:\n        #没有缓存，看有没有历史\n        try:\n            user_dict[user_id] = db_user_dict_list[user_id % 10][user_id]\n        except:\n            user_dict[user_id] = User()\n        return user_dict[user_id]\n</code></pre>",
      "rawMarkdown": "每个用户使用一个User 对象来动态更新状态特征。使用遍历的方法更新用户对象以及输出特征。\n用一个user_feat_dict 来存储所有用户和特征的关系。这是一个很大的字典。\n这个字典由两部分组成，\n\n第一部分。使用 sqlitedict 调用持久化在磁盘里的特征字典 （db_user_dict）\n第二部分。根据测试集的用户id 动态加载相应的用户特征到内存里。\n\n由于测试集中的用户很少。所以只提取测试集的用户id缓存到内存里。\n这样40G的线下特征文件，线上实际上内存占用约4G左右。\n```\n#用户字典\nuser_dict = {\n}\ndef get_user(user_id,user_dict):\n    # 如果缓存了\n    try:\n        return user_dict[user_id]\n    except:\n        #没有缓存，看有没有历史\n        try:\n            user_dict[user_id] = db_user_dict_list[user_id % 10][user_id]\n        except:\n            user_dict[user_id] = User()\n        return user_dict[user_id]\n```",
      "votes": null
    },
    {
      "id": "1142570",
      "postDate": "01/07/2021 13:27:07",
      "content": "<p>为什么不使用一个字典存储train当中每个用户起始的index，然后预测时，遇到一个用户从这个字典中搜索是否出现过该用户，出现过就使用你当前的方法仅对一个截取的起始index的train生成特征（只用一次），这样就不用上传存储一个40+G的字典文件，从我的实验来看，一个100+特征的lgb模型pipeline预测在3个小时内就可以完成。</p>",
      "rawMarkdown": "为什么不使用一个字典存储train当中每个用户起始的index，然后预测时，遇到一个用户从这个字典中搜索是否出现过该用户，出现过就使用你当前的方法仅对一个截取的起始index的train生成特征（只用一次），这样就不用上传存储一个40+G的字典文件，从我的实验来看，一个100+特征的lgb模型pipeline预测在3个小时内就可以完成。",
      "votes": null
    },
    {
      "id": "1142591",
      "postDate": "01/07/2021 13:40:37",
      "content": "<p>好主意，这里主要考虑到的是重新生成特征的时间。我存储预先生成好的特征。然后继续更新。速度会更快一些。上传40G的文件显然是很痛苦的，我也没这么做。所以我实际上是开了一个用于生成特征的私人kernel。这个kernel 跑完大约需要20min。然后提交的代码调用这个kernel。</p>",
      "rawMarkdown": "好主意，这里主要考虑到的是重新生成特征的时间。我存储预先生成好的特征。然后继续更新。速度会更快一些。上传40G的文件显然是很痛苦的，我也没这么做。所以我实际上是开了一个用于生成特征的私人kernel。这个kernel 跑完大约需要20min。然后提交的代码调用这个kernel。",
      "votes": null
    },
    {
      "id": "1142595",
      "postDate": "01/07/2021 13:43:45",
      "content": "<p>嗯，我觉得你目前遇到的提交问题可能就跟目前使用的db_user_dict相关，而且使用这种方法的人比较少，所以能给你参考建议的人也比较少，最后，祝好运。</p>",
      "rawMarkdown": "嗯，我觉得你目前遇到的提交问题可能就跟目前使用的db_user_dict相关，而且使用这种方法的人比较少，所以能给你参考建议的人也比较少，最后，祝好运。",
      "votes": null
    },
    {
      "id": "1142611",
      "postDate": "01/07/2021 13:54:13",
      "content": "<p>大佬，sqlite具体怎么样存嵌套字典的啊，找了好久也没见有可行的代码。。。我想把用户问题字典单纯给存到sqlite里面，怎么折腾都没成功。。。。</p>",
      "rawMarkdown": "大佬，sqlite具体怎么样存嵌套字典的啊，找了好久也没见有可行的代码。。。我想把用户问题字典单纯给存到sqlite里面，怎么折腾都没成功。。。。",
      "votes": null
    },
    {
      "id": "1142648",
      "postDate": "01/07/2021 14:13:19",
      "content": "<p>可能有关系，但是之前一直没有问题。所以这令我很头疼</p>",
      "rawMarkdown": "可能有关系，但是之前一直没有问题。所以这令我很头疼",
      "votes": null
    },
    {
      "id": "1142651",
      "postDate": "01/07/2021 14:16:13",
      "content": "<pre><code>def save_dict(df,i):\n   last_uid = None\n   user_obj = None\n   db_user_dict = SqliteDict('/kaggle/working/user_dict_{0}.sqlite'.format(i), autocommit=True)\n   for row in df.itertuples():\n       # 如果用户更新了，新建一个用户\n       if row.user_id != last_uid:\n           if last_uid is not None:\n               db_user_dict[last_uid] = user_obj\n           user_obj = User()\n       # 更新特征\n       user_obj.feed_sample(row)\n       last_uid = row.user_id\ndef build_dict(train_df):\n   n_jobs = 10\n   pool = multiprocessing.Pool(processes=5)\n\n   for i in range(5,10):\n       pool.apply_async(save_dict, (train_df[train_df.user_id % n_jobs == i],i,))\n   del train_df\n   pool.close()\n   pool.join()\nbuild_dict(train_df) \n</code></pre>",
      "rawMarkdown": "```\ndef save_dict(df,i):\n    last_uid = None\n    user_obj = None\n    db_user_dict = SqliteDict('/kaggle/working/user_dict_{0}.sqlite'.format(i), autocommit=True)\n    for row in df.itertuples():\n        # 如果用户更新了，新建一个用户\n        if row.user_id != last_uid:\n            if last_uid is not None:\n                db_user_dict[last_uid] = user_obj\n            user_obj = User()\n        # 更新特征\n        user_obj.feed_sample(row)\n        last_uid = row.user_id\n\ndef build_dict(train_df):\n    n_jobs = 10\n    pool = multiprocessing.Pool(processes=5)\n    \n    for i in range(5,10):\n        pool.apply_async(save_dict, (train_df[train_df.user_id % n_jobs == i],i,))\n    del train_df\n    pool.close()\n    pool.join()\n\nbuild_dict(train_df) \n```",
      "votes": null
    },
    {
      "id": "1142831",
      "postDate": "01/07/2021 16:05:57",
      "content": "<p>储存个题都这么费劲，都怪那些天天刷题的人，贡献这么多题。</p>",
      "rawMarkdown": "储存个题都这么费劲，都怪那些天天刷题的人，贡献这么多题。",
      "votes": null
    },
    {
      "id": "1142994",
      "postDate": "01/07/2021 17:58:37",
      "content": "<p>100+特征的LGB需要40G的字典文件，这都是些什么特征？我84个特征的LGB，用户特征pickle出来也就700多M，用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快</p>",
      "rawMarkdown": "100+特征的LGB需要40G的字典文件，这都是些什么特征？我84个特征的LGB，用户特征pickle出来也就700多M，用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快",
      "votes": null
    },
    {
      "id": "1143019",
      "postDate": "01/07/2021 18:07:26",
      "content": "<p>我觉得应该是因为存的数据是python的数据类型。不是numpy。</p>",
      "rawMarkdown": "我觉得应该是因为存的数据是python的数据类型。不是numpy。",
      "votes": null
    },
    {
      "id": "1143024",
      "postDate": "01/07/2021 18:10:27",
      "content": "<p>找到原因了。似乎官方把循环的第二个sample_prediction_df 变量给改掉了。之前我用这个变量提交的。</p>",
      "rawMarkdown": "找到原因了。似乎官方把循环的第二个sample_prediction_df 变量给改掉了。之前我用这个变量提交的。",
      "votes": null
    },
    {
      "id": "1143086",
      "postDate": "01/07/2021 18:42:44",
      "content": "<blockquote>\n  <p>我觉得应该是因为存的数据是python的数据类型。不是numpy。<br>\n  我也是存的python数据类型，用户级特征很多都是sparse的，没法存存成一个numpy的二维数组。如果你是用 <code>np.int32</code> 它和 int 一样都是占 28 byte</p>\n</blockquote>\n<pre><code>In [7]: sys.getsizeof(1)\nOut[7]: 28\n\nIn [8]: sys.getsizeof(np.int8(1))\nOut[8]: 25\n\nIn [9]: sys.getsizeof(np.int16(1))\nOut[9]: 26\n\nIn [10]: sys.getsizeof(np.int32(1))\nOut[10]: 28\n</code></pre>",
      "rawMarkdown": "> 我觉得应该是因为存的数据是python的数据类型。不是numpy。\n我也是存的python数据类型，用户级特征很多都是sparse的，没法存存成一个numpy的二维数组。如果你是用 `np.int32` 它和 int 一样都是占 28 byte\n```Python\nIn [7]: sys.getsizeof(1)\nOut[7]: 28\n\nIn [8]: sys.getsizeof(np.int8(1))\nOut[8]: 25\n\nIn [9]: sys.getsizeof(np.int16(1))\nOut[9]: 26\n\nIn [10]: sys.getsizeof(np.int32(1))\nOut[10]: 28\n```",
      "votes": null
    },
    {
      "id": "1143582",
      "postDate": "01/08/2021 00:43:13",
      "content": "<p>感谢你的分享，很有帮助，我一直被内存问题困扰，但看到你的-10我不厚道的笑了😅</p>",
      "rawMarkdown": "感谢你的分享，很有帮助，我一直被内存问题困扰，但看到你的-10我不厚道的笑了😅",
      "votes": null
    },
    {
      "id": "1143602",
      "postDate": "01/08/2021 00:58:04",
      "content": "<p>他们就硬卷呗哈哈哈。。。。。</p>",
      "rawMarkdown": "他们就硬卷呗哈哈哈。。。。。",
      "votes": null
    },
    {
      "id": "1143677",
      "postDate": "01/08/2021 02:03:58",
      "content": "<p>-16也太政治正确了…</p>",
      "rawMarkdown": "16也太政治正确了...",
      "votes": null
    },
    {
      "id": "1143687",
      "postDate": "01/08/2021 02:13:43",
      "content": "<p>最后会不会绝对值比正的最多的还多😐</p>",
      "rawMarkdown": "最后会不会绝对值比正的最多的还多😐",
      "votes": null
    },
    {
      "id": "1143691",
      "postDate": "01/08/2021 02:22:47",
      "content": "<p>为啥都给你-1？我给你点了个+1</p>",
      "rawMarkdown": "为啥都给你-1？我给你点了个+1",
      "votes": null
    },
    {
      "id": "1143713",
      "postDate": "01/08/2021 02:54:34",
      "content": "<p>🤥那就真的是 神贴留名了</p>",
      "rawMarkdown": "🤥那就真的是 神贴留名了",
      "votes": null
    },
    {
      "id": "1144077",
      "postDate": "01/08/2021 08:14:49",
      "content": "<p>这帮人不知道咋想的，我看到我读不懂的文字，最多无视，但不会踩一下</p>",
      "rawMarkdown": "这帮人不知道咋想的，我看到我读不懂的文字，最多无视，但不会踩一下",
      "votes": null
    },
    {
      "id": "1148057",
      "postDate": "01/10/2021 21:37:24",
      "content": "<p><a href=\"https://www.kaggle.com/infturing\" target=\"_blank\">@infturing</a> - I appreciate your sharing. But people giving downvotes have their reason. Kaggle is a competition + learning platform. And the competition rules mention that, the sharing should be public availabe (no private sharing). Despite you share on this forum which is public available, but using a language that only a minority of the participants can understand is not the idea of sharing that Kaggle want to have.</p>\n<p>If you really want to share your knowledge, please considering share it using English which most of us could understand. Let's not go to argue English vs. Chinese such political things, this is not the point. (Personally, I understand Chinese, so it doesn't bother me, but it would be better if you are willing to share in English). Let's focus on ML/DL that we all love.</p>\n<p>BTW, great job you have done!</p>",
      "rawMarkdown": "infturing - I appreciate your sharing. But people giving downvotes have their reason. Kaggle is a competition + learning platform. And the competition rules mention that, the sharing should be public availabe (no private sharing). Despite you share on this forum which is public available, but using a language that only a minority of the participants can understand is not the idea of sharing that Kaggle want to have.\n\nIf you really want to share your knowledge, please considering share it using English which most of us could understand. Let's not go to argue English vs. Chinese such political things, this is not the point. (Personally, I understand Chinese, so it doesn't bother me, but it would be better if you are willing to share in English). Let's focus on ML/DL that we all love.\n\nBTW, great job you have done!",
      "votes": null
    },
    {
      "id": "1148155",
      "postDate": "01/11/2021 00:07:39",
      "content": "<p>I'm not a \"杠精\". But according to your definition, sharing in English doesn't mean it's publicly available either. Lots of Chinese participants cannot understand English well. What they do is copy the English sentence to Google Translate to understand it. </p>\n<p>As long as the author shares it publicly, no matter what the language he uses, you can always get understood with the help of Google Translate. It's almost the same effort the author needs to make to write the article in English.</p>",
      "rawMarkdown": "I'm not a \"杠精\". But according to your definition, sharing in English doesn't mean it's publicly available either. Lots of Chinese participants cannot understand English well. What they do is copy the English sentence to Google Translate to understand it. \n\nAs long as the author shares it publicly, no matter what the language he uses, you can always get understood with the help of Google Translate. It's almost the same effort the author needs to make to write the article in English.",
      "votes": null
    },
    {
      "id": "1148455",
      "postDate": "01/11/2021 06:27:59",
      "content": "<p>Well,. I don't want to enter the debate further, just one final remark. I am surprised that a lot of chinese participants can't understand English even they are able to participate and take the challenge o Kaggle competitions ….</p>\n<p>Another remark is that, just think about if you are the host, and all the participants use their own language, Chinese, Japanese, French, etc ….. Would this situation is what you like?</p>\n<p>Anyway, you can make your choice, I am not here for arguing, and won't comments further. I would prefer to dive deeper in ML /  DL.</p>",
      "rawMarkdown": "Well,. I don't want to enter the debate further, just one final remark. I am surprised that a lot of chinese participants can't understand English even they are able to participate and take the challenge o Kaggle competitions ....\n\nAnother remark is that, just think about if you are the host, and all the participants use their own language, Chinese, Japanese, French, etc ..... Would this situation is what you like?\n\nAnyway, you can make your choice, I am not here for arguing, and won't comments further. I would prefer to dive deeper in ML /  DL.",
      "votes": null
    },
    {
      "id": "1148486",
      "postDate": "01/11/2021 07:09:00",
      "content": "<p>Just think about you were not familiar with English and to write your ideas in English you need to spend another 30 mins for translation. What do you think? I'm Chinese and all posts are in English because even I'm not good at English, writing in English is not a time-consuming thing for me. I don't want to enter the debate either, this is the last comment I leave under this thread. Bye and good luck!</p>",
      "rawMarkdown": "Just think about you were not familiar with English and to write your ideas in English you need to spend another 30 mins for translation. What do you think? I'm Chinese and all posts are in English because even I'm not good at English, writing in English is not a time-consuming thing for me. I don't want to enter the debate either, this is the last comment I leave under this thread. Bye and good luck!",
      "votes": null
    },
    {
      "id": "1148515",
      "postDate": "01/11/2021 07:31:26",
      "content": "<p>对啊，我之前还去日本的论坛蹭了蹭，google翻译超天下</p>",
      "rawMarkdown": "对啊，我之前还去日本的论坛蹭了蹭，google翻译超天下",
      "votes": null
    },
    {
      "id": "1148759",
      "postDate": "01/11/2021 11:16:57",
      "content": "<p>باهية برشا الفازة، مقصرتش في خدمتك يا طفل</p>",
      "rawMarkdown": "باهية برشا الفازة، مقصرتش في خدمتك يا طفل",
      "votes": null
    },
    {
      "id": "1149439",
      "postDate": "01/11/2021 20:46:46",
      "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> : about</p>\n<pre><code>用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快\n</code></pre>\n<p>Could you explain what are <code>用户级特征</code> and <code>依赖特征</code> you mentioned here, please? And maybe 1 or 2 examples for each. Thank you in advance.</p>",
      "rawMarkdown": "wuwenmin : about\n\n```\n用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快\n```\n\nCould you explain what are `用户级特征` and `依赖特征` you mentioned here, please? And maybe 1 or 2 examples for each. Thank you in advance.",
      "votes": null
    },
    {
      "id": "1149514",
      "postDate": "01/11/2021 22:40:39",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>  For all the features that I used in the final submission can refer to <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596#1143616\" target=\"_blank\">[0.805 Private, 42nd place] LGB + 5 RAINT+ Ensemble</a><br>\nLet me use <code>user_ans_cnt</code>, <code>user_ans_corr_cnt</code>, and <code>user_ans_acc</code> as an example. I don't need to store the <code>user_ans_acc</code> for each user, since <code>user_ans_acc</code> = <code>user_ans_corr_cnt</code> / <code>user_ans_cnt</code></p>\n<p>In my implementation, I further optimize this part. For each <code>test_df</code> I extract non-dependent features(e.g. <code>user_ans_cnt</code>) first to get a feature array <code>arr (MxN)</code>, where <code>M=test_df.shape[0]</code> and <code>N</code> is the number of features. Then all dependent features can be calculated with:<br>\n<code>arr[dependent_fea_idx] = cal_func(*[arr[idx_] for idx_ in dep_idx])</code></p>\n<p>All these indices and calculation functions are initialized at the beginning. During the inference period, they're just a list of extraction functions (no <code>dict</code>, no <code>if/else</code>) here. Refer to <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529\" target=\"_blank\">Tricks &amp; Tips to Dramatically Speed Up Submission Running</a></p>",
      "rawMarkdown": "Hi @yihdarshieh  For all the features that I used in the final submission can refer to [[0.805 Private, 42nd place] LGB + 5 RAINT+ Ensemble](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596#1143616)\nLet me use `user_ans_cnt`, `user_ans_corr_cnt`, and `user_ans_acc` as an example. I don't need to store the `user_ans_acc` for each user, since `user_ans_acc` = `user_ans_corr_cnt` / `user_ans_cnt`\n\nIn my implementation, I further optimize this part. For each `test_df` I extract non-dependent features(e.g. `user_ans_cnt`) first to get a feature array `arr (MxN)`, where `M=test_df.shape[0]` and `N` is the number of features. Then all dependent features can be calculated with:\n`arr[dependent_fea_idx] = cal_func(*[arr[idx_] for idx_ in dep_idx])`\n\nAll these indices and calculation functions are initialized at the beginning. During the inference period, they're just a list of extraction functions (no `dict`, no `if/else`) here. Refer to [Tricks & Tips to Dramatically Speed Up Submission Running](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529)",
      "votes": null
    },
    {
      "id": "1149523",
      "postDate": "01/11/2021 22:52:49",
      "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> Thank you for the detailed explanation. I need a bit time to fully understand, but let me ask one more question to be sure. user_ans_corr_cnt is also a non dependent features in your definition above right? And the optimal<br>\n way of calculation is applied to the calculation of  answer correctness accuracy from the 2 counts, in the above example, Right?</p>",
      "rawMarkdown": "wuwenmin Thank you for the detailed explanation. I need a bit time to fully understand, but let me ask one more question to be sure. user_ans_corr_cnt is also a non dependent features in your definition above right? And the optimal\n way of calculation is applied to the calculation of  answer correctness accuracy from the 2 counts, in the above example, Right?",
      "votes": null
    },
    {
      "id": "1149548",
      "postDate": "01/11/2021 23:47:25",
      "content": "<p>Yup, correct</p>",
      "rawMarkdown": "Yup, correct",
      "votes": null
    },
    {
      "id": "1152708",
      "postDate": "01/14/2021 11:38:18",
      "content": "<p>感谢分享，线上好像看不到内存吧， 4G 是在本地模拟测试的吗？</p>",
      "rawMarkdown": "感谢分享，线上好像看不到内存吧， 4G 是在本地模拟测试的吗？",
      "votes": null
    },
    {
      "id": "1152738",
      "postDate": "01/14/2021 12:14:14",
      "content": "<p>You can see the RAM usage by clicking the button near 'Draft session' on the top-right corner of the editor</p>\n<p><a href=\"https://ibb.co/52JnXs1\"><img src=\"https://i.ibb.co/yVxYmgN/Capture333.png\" alt=\"Capture333\"></a></p>",
      "rawMarkdown": "You can see the RAM usage by clicking the button near 'Draft session' on the top-right corner of the editor\n\n<a href=\"https://ibb.co/52JnXs1\"><img src=\"https://i.ibb.co/yVxYmgN/Capture333.png\" alt=\"Capture333\" border=\"0\" /></a>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1142570,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "01/07/2021 13:27:07",
      "content": "<p>为什么不使用一个字典存储train当中每个用户起始的index，然后预测时，遇到一个用户从这个字典中搜索是否出现过该用户，出现过就使用你当前的方法仅对一个截取的起始index的train生成特征（只用一次），这样就不用上传存储一个40+G的字典文件，从我的实验来看，一个100+特征的lgb模型pipeline预测在3个小时内就可以完成。</p>",
      "votes": null,
      "replies": [
        {
          "id": 1142591,
          "author_name": "infturing",
          "author_url": "",
          "post_date": "01/07/2021 13:40:37",
          "content": "<p>好主意，这里主要考虑到的是重新生成特征的时间。我存储预先生成好的特征。然后继续更新。速度会更快一些。上传40G的文件显然是很痛苦的，我也没这么做。所以我实际上是开了一个用于生成特征的私人kernel。这个kernel 跑完大约需要20min。然后提交的代码调用这个kernel。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142595,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "01/07/2021 13:43:45",
          "content": "<p>嗯，我觉得你目前遇到的提交问题可能就跟目前使用的db_user_dict相关，而且使用这种方法的人比较少，所以能给你参考建议的人也比较少，最后，祝好运。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142648,
          "author_name": "infturing",
          "author_url": "",
          "post_date": "01/07/2021 14:13:19",
          "content": "<p>可能有关系，但是之前一直没有问题。所以这令我很头疼</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142994,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/07/2021 17:58:37",
          "content": "<p>100+特征的LGB需要40G的字典文件，这都是些什么特征？我84个特征的LGB，用户特征pickle出来也就700多M，用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143019,
          "author_name": "infturing",
          "author_url": "",
          "post_date": "01/07/2021 18:07:26",
          "content": "<p>我觉得应该是因为存的数据是python的数据类型。不是numpy。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143024,
          "author_name": "infturing",
          "author_url": "",
          "post_date": "01/07/2021 18:10:27",
          "content": "<p>找到原因了。似乎官方把循环的第二个sample_prediction_df 变量给改掉了。之前我用这个变量提交的。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143086,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/07/2021 18:42:44",
          "content": "<blockquote>\n  <p>我觉得应该是因为存的数据是python的数据类型。不是numpy。<br>\n  我也是存的python数据类型，用户级特征很多都是sparse的，没法存存成一个numpy的二维数组。如果你是用 <code>np.int32</code> 它和 int 一样都是占 28 byte</p>\n</blockquote>\n<pre><code>In [7]: sys.getsizeof(1)\nOut[7]: 28\n\nIn [8]: sys.getsizeof(np.int8(1))\nOut[8]: 25\n\nIn [9]: sys.getsizeof(np.int16(1))\nOut[9]: 26\n\nIn [10]: sys.getsizeof(np.int32(1))\nOut[10]: 28\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149439,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/11/2021 20:46:46",
          "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> : about</p>\n<pre><code>用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快\n</code></pre>\n<p>Could you explain what are <code>用户级特征</code> and <code>依赖特征</code> you mentioned here, please? And maybe 1 or 2 examples for each. Thank you in advance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149514,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/11/2021 22:40:39",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>  For all the features that I used in the final submission can refer to <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596#1143616\" target=\"_blank\">[0.805 Private, 42nd place] LGB + 5 RAINT+ Ensemble</a><br>\nLet me use <code>user_ans_cnt</code>, <code>user_ans_corr_cnt</code>, and <code>user_ans_acc</code> as an example. I don't need to store the <code>user_ans_acc</code> for each user, since <code>user_ans_acc</code> = <code>user_ans_corr_cnt</code> / <code>user_ans_cnt</code></p>\n<p>In my implementation, I further optimize this part. For each <code>test_df</code> I extract non-dependent features(e.g. <code>user_ans_cnt</code>) first to get a feature array <code>arr (MxN)</code>, where <code>M=test_df.shape[0]</code> and <code>N</code> is the number of features. Then all dependent features can be calculated with:<br>\n<code>arr[dependent_fea_idx] = cal_func(*[arr[idx_] for idx_ in dep_idx])</code></p>\n<p>All these indices and calculation functions are initialized at the beginning. During the inference period, they're just a list of extraction functions (no <code>dict</code>, no <code>if/else</code>) here. Refer to <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529\" target=\"_blank\">Tricks &amp; Tips to Dramatically Speed Up Submission Running</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149523,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/11/2021 22:52:49",
          "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> Thank you for the detailed explanation. I need a bit time to fully understand, but let me ask one more question to be sure. user_ans_corr_cnt is also a non dependent features in your definition above right? And the optimal<br>\n way of calculation is applied to the calculation of  answer correctness accuracy from the 2 counts, in the above example, Right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149548,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/11/2021 23:47:25",
          "content": "<p>Yup, correct</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1142611,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "01/07/2021 13:54:13",
      "content": "<p>大佬，sqlite具体怎么样存嵌套字典的啊，找了好久也没见有可行的代码。。。我想把用户问题字典单纯给存到sqlite里面，怎么折腾都没成功。。。。</p>",
      "votes": null,
      "replies": [
        {
          "id": 1142651,
          "author_name": "infturing",
          "author_url": "",
          "post_date": "01/07/2021 14:16:13",
          "content": "<pre><code>def save_dict(df,i):\n   last_uid = None\n   user_obj = None\n   db_user_dict = SqliteDict('/kaggle/working/user_dict_{0}.sqlite'.format(i), autocommit=True)\n   for row in df.itertuples():\n       # 如果用户更新了，新建一个用户\n       if row.user_id != last_uid:\n           if last_uid is not None:\n               db_user_dict[last_uid] = user_obj\n           user_obj = User()\n       # 更新特征\n       user_obj.feed_sample(row)\n       last_uid = row.user_id\ndef build_dict(train_df):\n   n_jobs = 10\n   pool = multiprocessing.Pool(processes=5)\n\n   for i in range(5,10):\n       pool.apply_async(save_dict, (train_df[train_df.user_id % n_jobs == i],i,))\n   del train_df\n   pool.close()\n   pool.join()\nbuild_dict(train_df) \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142831,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "01/07/2021 16:05:57",
          "content": "<p>储存个题都这么费劲，都怪那些天天刷题的人，贡献这么多题。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143602,
          "author_name": "qq1623620766",
          "author_url": "",
          "post_date": "01/08/2021 00:58:04",
          "content": "<p>他们就硬卷呗哈哈哈。。。。。</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143582,
      "author_name": "louieshao",
      "author_url": "",
      "post_date": "01/08/2021 00:43:13",
      "content": "<p>感谢你的分享，很有帮助，我一直被内存问题困扰，但看到你的-10我不厚道的笑了😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144077,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/08/2021 08:14:49",
          "content": "<p>这帮人不知道咋想的，我看到我读不懂的文字，最多无视，但不会踩一下</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148515,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "01/11/2021 07:31:26",
          "content": "<p>对啊，我之前还去日本的论坛蹭了蹭，google翻译超天下</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143677,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "01/08/2021 02:03:58",
      "content": "<p>-16也太政治正确了…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143687,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "01/08/2021 02:13:43",
          "content": "<p>最后会不会绝对值比正的最多的还多😐</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143713,
          "author_name": "xiaowangiiiii",
          "author_url": "",
          "post_date": "01/08/2021 02:54:34",
          "content": "<p>🤥那就真的是 神贴留名了</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143691,
      "author_name": "gdwainwnog",
      "author_url": "",
      "post_date": "01/08/2021 02:22:47",
      "content": "<p>为啥都给你-1？我给你点了个+1</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148057,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/10/2021 21:37:24",
      "content": "<p><a href=\"https://www.kaggle.com/infturing\" target=\"_blank\">@infturing</a> - I appreciate your sharing. But people giving downvotes have their reason. Kaggle is a competition + learning platform. And the competition rules mention that, the sharing should be public availabe (no private sharing). Despite you share on this forum which is public available, but using a language that only a minority of the participants can understand is not the idea of sharing that Kaggle want to have.</p>\n<p>If you really want to share your knowledge, please considering share it using English which most of us could understand. Let's not go to argue English vs. Chinese such political things, this is not the point. (Personally, I understand Chinese, so it doesn't bother me, but it would be better if you are willing to share in English). Let's focus on ML/DL that we all love.</p>\n<p>BTW, great job you have done!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148155,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/11/2021 00:07:39",
          "content": "<p>I'm not a \"杠精\". But according to your definition, sharing in English doesn't mean it's publicly available either. Lots of Chinese participants cannot understand English well. What they do is copy the English sentence to Google Translate to understand it. </p>\n<p>As long as the author shares it publicly, no matter what the language he uses, you can always get understood with the help of Google Translate. It's almost the same effort the author needs to make to write the article in English.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148455,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/11/2021 06:27:59",
          "content": "<p>Well,. I don't want to enter the debate further, just one final remark. I am surprised that a lot of chinese participants can't understand English even they are able to participate and take the challenge o Kaggle competitions ….</p>\n<p>Another remark is that, just think about if you are the host, and all the participants use their own language, Chinese, Japanese, French, etc ….. Would this situation is what you like?</p>\n<p>Anyway, you can make your choice, I am not here for arguing, and won't comments further. I would prefer to dive deeper in ML /  DL.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148486,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/11/2021 07:09:00",
          "content": "<p>Just think about you were not familiar with English and to write your ideas in English you need to spend another 30 mins for translation. What do you think? I'm Chinese and all posts are in English because even I'm not good at English, writing in English is not a time-consuming thing for me. I don't want to enter the debate either, this is the last comment I leave under this thread. Bye and good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148759,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "01/11/2021 11:16:57",
      "content": "<p>باهية برشا الفازة، مقصرتش في خدمتك يا طفل</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1152708,
      "author_name": "xjz98k",
      "author_url": "",
      "post_date": "01/14/2021 11:38:18",
      "content": "<p>感谢分享，线上好像看不到内存吧， 4G 是在本地模拟测试的吗？</p>",
      "votes": null,
      "replies": [
        {
          "id": 1152738,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/14/2021 12:14:14",
          "content": "<p>You can see the RAM usage by clicking the button near 'Draft session' on the top-right corner of the editor</p>\n<p><a href=\"https://ibb.co/52JnXs1\"><img src=\"https://i.ibb.co/yVxYmgN/Capture333.png\" alt=\"Capture333\"></a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1142560": "每个用户使用一个User 对象来动态更新状态特征。使用遍历的方法更新用户对象以及输出特征。\n用一个user_feat_dict 来存储所有用户和特征的关系。这是一个很大的字典。\n这个字典由两部分组成，\n\n第一部分。使用 sqlitedict 调用持久化在磁盘里的特征字典 （db_user_dict）\n第二部分。根据测试集的用户id 动态加载相应的用户特征到内存里。\n\n由于测试集中的用户很少。所以只提取测试集的用户id缓存到内存里。\n这样40G的线下特征文件，线上实际上内存占用约4G左右。\n```\n#用户字典\nuser_dict = {\n}\ndef get_user(user_id,user_dict):\n    # 如果缓存了\n    try:\n        return user_dict[user_id]\n    except:\n        #没有缓存，看有没有历史\n        try:\n            user_dict[user_id] = db_user_dict_list[user_id % 10][user_id]\n        except:\n            user_dict[user_id] = User()\n        return user_dict[user_id]\n```",
    "1142570": "为什么不使用一个字典存储train当中每个用户起始的index，然后预测时，遇到一个用户从这个字典中搜索是否出现过该用户，出现过就使用你当前的方法仅对一个截取的起始index的train生成特征（只用一次），这样就不用上传存储一个40+G的字典文件，从我的实验来看，一个100+特征的lgb模型pipeline预测在3个小时内就可以完成。",
    "1142591": "好主意，这里主要考虑到的是重新生成特征的时间。我存储预先生成好的特征。然后继续更新。速度会更快一些。上传40G的文件显然是很痛苦的，我也没这么做。所以我实际上是开了一个用于生成特征的私人kernel。这个kernel 跑完大约需要20min。然后提交的代码调用这个kernel。",
    "1142595": "嗯，我觉得你目前遇到的提交问题可能就跟目前使用的db_user_dict相关，而且使用这种方法的人比较少，所以能给你参考建议的人也比较少，最后，祝好运。",
    "1142611": "大佬，sqlite具体怎么样存嵌套字典的啊，找了好久也没见有可行的代码。。。我想把用户问题字典单纯给存到sqlite里面，怎么折腾都没成功。。。。",
    "1142648": "可能有关系，但是之前一直没有问题。所以这令我很头疼",
    "1142651": "```\ndef save_dict(df,i):\n    last_uid = None\n    user_obj = None\n    db_user_dict = SqliteDict('/kaggle/working/user_dict_{0}.sqlite'.format(i), autocommit=True)\n    for row in df.itertuples():\n        # 如果用户更新了，新建一个用户\n        if row.user_id != last_uid:\n            if last_uid is not None:\n                db_user_dict[last_uid] = user_obj\n            user_obj = User()\n        # 更新特征\n        user_obj.feed_sample(row)\n        last_uid = row.user_id\n\ndef build_dict(train_df):\n    n_jobs = 10\n    pool = multiprocessing.Pool(processes=5)\n    \n    for i in range(5,10):\n        pool.apply_async(save_dict, (train_df[train_df.user_id % n_jobs == i],i,))\n    del train_df\n    pool.close()\n    pool.join()\n\nbuild_dict(train_df) \n```",
    "1142831": "储存个题都这么费劲，都怪那些天天刷题的人，贡献这么多题。",
    "1142994": "100+特征的LGB需要40G的字典文件，这都是些什么特征？我84个特征的LGB，用户特征pickle出来也就700多M，用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快",
    "1143019": "我觉得应该是因为存的数据是python的数据类型。不是numpy。",
    "1143024": "找到原因了。似乎官方把循环的第二个sample_prediction_df 变量给改掉了。之前我用这个变量提交的。",
    "1143086": "> 我觉得应该是因为存的数据是python的数据类型。不是numpy。\n我也是存的python数据类型，用户级特征很多都是sparse的，没法存存成一个numpy的二维数组。如果你是用 `np.int32` 它和 int 一样都是占 28 byte\n```Python\nIn [7]: sys.getsizeof(1)\nOut[7]: 28\n\nIn [8]: sys.getsizeof(np.int8(1))\nOut[8]: 25\n\nIn [9]: sys.getsizeof(np.int16(1))\nOut[9]: 26\n\nIn [10]: sys.getsizeof(np.int32(1))\nOut[10]: 28\n```",
    "1143582": "感谢你的分享，很有帮助，我一直被内存问题困扰，但看到你的-10我不厚道的笑了😅",
    "1143602": "他们就硬卷呗哈哈哈。。。。。",
    "1143677": "16也太政治正确了...",
    "1143687": "最后会不会绝对值比正的最多的还多😐",
    "1143691": "为啥都给你-1？我给你点了个+1",
    "1143713": "🤥那就真的是 神贴留名了",
    "1144077": "这帮人不知道咋想的，我看到我读不懂的文字，最多无视，但不会踩一下",
    "1148057": "infturing - I appreciate your sharing. But people giving downvotes have their reason. Kaggle is a competition + learning platform. And the competition rules mention that, the sharing should be public availabe (no private sharing). Despite you share on this forum which is public available, but using a language that only a minority of the participants can understand is not the idea of sharing that Kaggle want to have.\n\nIf you really want to share your knowledge, please considering share it using English which most of us could understand. Let's not go to argue English vs. Chinese such political things, this is not the point. (Personally, I understand Chinese, so it doesn't bother me, but it would be better if you are willing to share in English). Let's focus on ML/DL that we all love.\n\nBTW, great job you have done!",
    "1148155": "I'm not a \"杠精\". But according to your definition, sharing in English doesn't mean it's publicly available either. Lots of Chinese participants cannot understand English well. What they do is copy the English sentence to Google Translate to understand it. \n\nAs long as the author shares it publicly, no matter what the language he uses, you can always get understood with the help of Google Translate. It's almost the same effort the author needs to make to write the article in English.",
    "1148455": "Well,. I don't want to enter the debate further, just one final remark. I am surprised that a lot of chinese participants can't understand English even they are able to participate and take the challenge o Kaggle competitions ....\n\nAnother remark is that, just think about if you are the host, and all the participants use their own language, Chinese, Japanese, French, etc ..... Would this situation is what you like?\n\nAnyway, you can make your choice, I am not here for arguing, and won't comments further. I would prefer to dive deeper in ML /  DL.",
    "1148486": "Just think about you were not familiar with English and to write your ideas in English you need to spend another 30 mins for translation. What do you think? I'm Chinese and all posts are in English because even I'm not good at English, writing in English is not a time-consuming thing for me. I don't want to enter the debate either, this is the last comment I leave under this thread. Bye and good luck!",
    "1148515": "对啊，我之前还去日本的论坛蹭了蹭，google翻译超天下",
    "1148759": "باهية برشا الفازة، مقصرتش في خدمتك يا طفل",
    "1149439": "wuwenmin : about\n\n```\n用户级特征只需存不同纬度的count，依赖特征实时算比查字典更快\n```\n\nCould you explain what are `用户级特征` and `依赖特征` you mentioned here, please? And maybe 1 or 2 examples for each. Thank you in advance.",
    "1149514": "Hi @yihdarshieh  For all the features that I used in the final submission can refer to [[0.805 Private, 42nd place] LGB + 5 RAINT+ Ensemble](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596#1143616)\nLet me use `user_ans_cnt`, `user_ans_corr_cnt`, and `user_ans_acc` as an example. I don't need to store the `user_ans_acc` for each user, since `user_ans_acc` = `user_ans_corr_cnt` / `user_ans_cnt`\n\nIn my implementation, I further optimize this part. For each `test_df` I extract non-dependent features(e.g. `user_ans_cnt`) first to get a feature array `arr (MxN)`, where `M=test_df.shape[0]` and `N` is the number of features. Then all dependent features can be calculated with:\n`arr[dependent_fea_idx] = cal_func(*[arr[idx_] for idx_ in dep_idx])`\n\nAll these indices and calculation functions are initialized at the beginning. During the inference period, they're just a list of extraction functions (no `dict`, no `if/else`) here. Refer to [Tricks & Tips to Dramatically Speed Up Submission Running](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529)",
    "1149523": "wuwenmin Thank you for the detailed explanation. I need a bit time to fully understand, but let me ask one more question to be sure. user_ans_corr_cnt is also a non dependent features in your definition above right? And the optimal\n way of calculation is applied to the calculation of  answer correctness accuracy from the 2 counts, in the above example, Right?",
    "1149548": "Yup, correct",
    "1152708": "感谢分享，线上好像看不到内存吧， 4G 是在本地模拟测试的吗？",
    "1152738": "You can see the RAM usage by clicking the button near 'Draft session' on the top-right corner of the editor\n\n<a href=\"https://ibb.co/52JnXs1\"><img src=\"https://i.ibb.co/yVxYmgN/Capture333.png\" alt=\"Capture333\" border=\"0\" /></a>"
  },
  "source": "meta"
}