{
  "id": 56105,
  "title": "How to use more features to training with dark numpy magic...",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56105",
  "author_name": "",
  "post_date": "2018-05-06T06:04:37.538905900Z",
  "votes": 87,
  "comment_count": 12,
  "views": 0,
  "content": "<p>This competition is very hard according to large data, so all kaggler's try to use some tricks to train with more data in memory constraint environment. Thanks to all who shared their approaches - it helps a lot! Here I want to add one more trick to bag of community solution - how to combine several files with group of pickled attrs to one train/val without memory spikes.  </p>\n\n<p>Let we have separate group of attrs pickled in different files 'attrs_?????' in directory './data/'. The each file is concatenated [train,test] part of train dataframe.</p>\n\n<pre><code>split_train = 11111111     # border between train/test in pickle (length of our train)\nstart_skip = 0   # if we want strip our train\nn_attrs = 25  # as is\n\npickle_list = ['attrs_xxxx',   'attrs_base_cnt',   'attrs_nunique']   # list of pickled attrs\ndel_cols = [  'click_time',  'day', ]  # attrs to be deleted in final train\n</code></pre>\n\n<p>At first we create memmap file where we will store our combined train dataframe</p>\n\n<pre><code>si = 0   \nmmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='w+', shape=(split_train, n_attrs))\n</code></pre>\n\n<p>define helper function</p>\n\n<pre><code>def load_attrs(fname, data_dir='./data/'):\n    fname = data_dir + fname + '.pkl'\n    print('loading {}... '.format(fname))\n    return pd.read_pickle(fname)\n</code></pre>\n\n<p>and convert our parts (of attrs) to appropriate position in the memmap file. Note - we have to save columns of our combined dataframe as memmap is just numpy array (less or more :) ) and it doesn't have information about columns.</p>\n\n<pre><code>columns = []\nfor pkl in pickle_list:  \n    _temp = load_attrs(pkl)\n\n    _columns  = [x for x in _temp.columns if x not in del_cols]  \n    columns = columns + _columns\n\n    nodel_ind = [_temp.columns.tolist().index(x) for x in _temp.columns if x not in del_cols]  \n\n    _temp = _temp.iloc[start_skip:split_train, nodel_ind]  \n\n    ei = _temp.values.shape[1]  \n    mmap[:, si:si+ei] = _temp.values  \n    si += ei  \n\n    del _temp  \n    gc.collect()\n</code></pre>\n\n<p>then we flushed our memmap file and close it.</p>\n\n<pre><code>mmap.flush()\ndel mmap\ngc.collect()\n</code></pre>\n\n<p>reopen our memmap in 'read-only' mode and extract train/val as numpy array.</p>\n\n<pre><code>mmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='r', shape=(split_train, n_attrs))\n_train = np.array(mmap[start_skip:-val_size])\n_val = np.array(mmap[-val_size:])\n</code></pre>\n\n<p>get y</p>\n\n<pre><code>_train = _train[:, columns.index('is_attributed')]\n_val = _val[:, columns.index('is_attributed')]\n</code></pre>\n\n<p>prepare set of columns to use in training and get indexes of them.</p>\n\n<pre><code>use_columns = [ 'app',  'channel',  'device',  'os',  'hour', 'ip_cnt']\nmuse_columns = [columns.index(x) for x in use_columns]\n</code></pre>\n\n<p>prepare lgb.Datasets</p>\n\n<pre><code>d_train = _train[:, muse_columns]\nxgtrain = lgb.Dataset(d_train, label=yy_train, **dataset_params)\nd_val = mXX_val[:, muse_columns]\nxgvalid = lgb.Dataset(d_val, label=yy_val, reference=xgtrain, **dataset_params)\n</code></pre>\n\n<p>train</p>\n\n<pre><code>bst = lgb.train(model_params, xgtrain, valid_sets=[xgvalid], valid_names=['valid'], evals_result=evals_results, **fit_params)\n</code></pre>\n\n<p>and voila! No spikes, no memory leaking, no ARGHHHH ) <br>\nHope this hepls.</p>\n\n<p>Happy kaggling! <br>\nKruegger</p>\n\n<p>P.S. you can upvote if you find this useful )</p>",
  "messages": [
    {
      "id": "323767",
      "postDate": "05/06/2018 06:04:37",
      "content": "<p>This competition is very hard according to large data, so all kaggler's try to use some tricks to train with more data in memory constraint environment. Thanks to all who shared their approaches - it helps a lot! Here I want to add one more trick to bag of community solution - how to combine several files with group of pickled attrs to one train/val without memory spikes.  </p>\n\n<p>Let we have separate group of attrs pickled in different files 'attrs_?????' in directory './data/'. The each file is concatenated [train,test] part of train dataframe.</p>\n\n<pre><code>split_train = 11111111     # border between train/test in pickle (length of our train)\nstart_skip = 0   # if we want strip our train\nn_attrs = 25  # as is\n\npickle_list = ['attrs_xxxx',   'attrs_base_cnt',   'attrs_nunique']   # list of pickled attrs\ndel_cols = [  'click_time',  'day', ]  # attrs to be deleted in final train\n</code></pre>\n\n<p>At first we create memmap file where we will store our combined train dataframe</p>\n\n<pre><code>si = 0   \nmmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='w+', shape=(split_train, n_attrs))\n</code></pre>\n\n<p>define helper function</p>\n\n<pre><code>def load_attrs(fname, data_dir='./data/'):\n    fname = data_dir + fname + '.pkl'\n    print('loading {}... '.format(fname))\n    return pd.read_pickle(fname)\n</code></pre>\n\n<p>and convert our parts (of attrs) to appropriate position in the memmap file. Note - we have to save columns of our combined dataframe as memmap is just numpy array (less or more :) ) and it doesn't have information about columns.</p>\n\n<pre><code>columns = []\nfor pkl in pickle_list:  \n    _temp = load_attrs(pkl)\n\n    _columns  = [x for x in _temp.columns if x not in del_cols]  \n    columns = columns + _columns\n\n    nodel_ind = [_temp.columns.tolist().index(x) for x in _temp.columns if x not in del_cols]  \n\n    _temp = _temp.iloc[start_skip:split_train, nodel_ind]  \n\n    ei = _temp.values.shape[1]  \n    mmap[:, si:si+ei] = _temp.values  \n    si += ei  \n\n    del _temp  \n    gc.collect()\n</code></pre>\n\n<p>then we flushed our memmap file and close it.</p>\n\n<pre><code>mmap.flush()\ndel mmap\ngc.collect()\n</code></pre>\n\n<p>reopen our memmap in 'read-only' mode and extract train/val as numpy array.</p>\n\n<pre><code>mmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='r', shape=(split_train, n_attrs))\n_train = np.array(mmap[start_skip:-val_size])\n_val = np.array(mmap[-val_size:])\n</code></pre>\n\n<p>get y</p>\n\n<pre><code>_train = _train[:, columns.index('is_attributed')]\n_val = _val[:, columns.index('is_attributed')]\n</code></pre>\n\n<p>prepare set of columns to use in training and get indexes of them.</p>\n\n<pre><code>use_columns = [ 'app',  'channel',  'device',  'os',  'hour', 'ip_cnt']\nmuse_columns = [columns.index(x) for x in use_columns]\n</code></pre>\n\n<p>prepare lgb.Datasets</p>\n\n<pre><code>d_train = _train[:, muse_columns]\nxgtrain = lgb.Dataset(d_train, label=yy_train, **dataset_params)\nd_val = mXX_val[:, muse_columns]\nxgvalid = lgb.Dataset(d_val, label=yy_val, reference=xgtrain, **dataset_params)\n</code></pre>\n\n<p>train</p>\n\n<pre><code>bst = lgb.train(model_params, xgtrain, valid_sets=[xgvalid], valid_names=['valid'], evals_result=evals_results, **fit_params)\n</code></pre>\n\n<p>and voila! No spikes, no memory leaking, no ARGHHHH ) <br>\nHope this hepls.</p>\n\n<p>Happy kaggling! <br>\nKruegger</p>\n\n<p>P.S. you can upvote if you find this useful )</p>",
      "rawMarkdown": "This competition is very hard according to large data, so all kaggler's try to use some tricks to train with more data in memory constraint environment. Thanks to all who shared their approaches - it helps a lot! Here I want to add one more trick to bag of community solution - how to combine several files with group of pickled attrs to one train/val without memory spikes.  \n\nLet we have separate group of attrs pickled in different files 'attrs_?????' in directory './data/'. The each file is concatenated [train,test] part of train dataframe.\n\n    split_train = 11111111     # border between train/test in pickle (length of our train)\n    start_skip = 0   # if we want strip our train\n    n_attrs = 25  # as is\n    \n    pickle_list = ['attrs_xxxx',   'attrs_base_cnt',   'attrs_nunique']   # list of pickled attrs\n    del_cols = [  'click_time',  'day', ]  # attrs to be deleted in final train\n\nAt first we create memmap file where we will store our combined train dataframe\n\n    si = 0   \n    mmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='w+', shape=(split_train, n_attrs))\n\ndefine helper function\n\n    def load_attrs(fname, data_dir='./data/'):\n        fname = data_dir + fname + '.pkl'\n        print('loading {}... '.format(fname))\n        return pd.read_pickle(fname)\n\nand convert our parts (of attrs) to appropriate position in the memmap file. Note - we have to save columns of our combined dataframe as memmap is just numpy array (less or more :) ) and it doesn't have information about columns.\n\n    columns = []\n    for pkl in pickle_list:  \n        _temp = load_attrs(pkl)\n    \n        _columns  = [x for x in _temp.columns if x not in del_cols]  \n        columns = columns + _columns\n    \n        nodel_ind = [_temp.columns.tolist().index(x) for x in _temp.columns if x not in del_cols]  \n        \n        _temp = _temp.iloc[start_skip:split_train, nodel_ind]  \n        \n        ei = _temp.values.shape[1]  \n        mmap[:, si:si+ei] = _temp.values  \n        si += ei  \n         \n        del _temp  \n        gc.collect()\n\nthen we flushed our memmap file and close it.\n\n    mmap.flush()\n    del mmap\n    gc.collect()\n\n\nreopen our memmap in 'read-only' mode and extract train/val as numpy array.\n\n    mmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='r', shape=(split_train, n_attrs))\n    _train = np.array(mmap[start_skip:-val_size])\n    _val = np.array(mmap[-val_size:])\n\nget y\n\n    _train = _train[:, columns.index('is_attributed')]\n    _val = _val[:, columns.index('is_attributed')]\n\nprepare set of columns to use in training and get indexes of them.\n\n    use_columns = [ 'app',  'channel',  'device',  'os',  'hour', 'ip_cnt']\n    muse_columns = [columns.index(x) for x in use_columns]\n\nprepare lgb.Datasets\n\n    d_train = _train[:, muse_columns]\n    xgtrain = lgb.Dataset(d_train, label=yy_train, **dataset_params)\n    d_val = mXX_val[:, muse_columns]\n    xgvalid = lgb.Dataset(d_val, label=yy_val, reference=xgtrain, **dataset_params)\n\ntrain\n\n\n    bst = lgb.train(model_params, xgtrain, valid_sets=[xgvalid], valid_names=['valid'], evals_result=evals_results, **fit_params)\n\nand voila! No spikes, no memory leaking, no ARGHHHH )  \nHope this hepls.\n\nHappy kaggling!  \nKruegger\n\nP.S. you can upvote if you find this useful )",
      "votes": null
    },
    {
      "id": "323802",
      "postDate": "05/06/2018 09:17:03",
      "content": "<p>Hi, thanks for sharing.  Isn't this more or less equivalent to the two_stage of lgb?  Or it is better?</p>",
      "rawMarkdown": "Hi, thanks for sharing.  Isn't this more or less equivalent to the two_stage of lgb?  Or it is better?",
      "votes": null
    },
    {
      "id": "323832",
      "postDate": "05/06/2018 11:10:24",
      "content": "<p>This would be a useful technique. </p>",
      "rawMarkdown": "This would be a useful technique.",
      "votes": null
    },
    {
      "id": "323833",
      "postDate": "05/06/2018 11:11:18",
      "content": "<p>Does the two round parameter works for LGB python package or it’s just for LGB command line?</p>",
      "rawMarkdown": "Does the two round parameter works for LGB python package or it’s just for LGB command line?",
      "votes": null
    },
    {
      "id": "323838",
      "postDate": "05/06/2018 11:20:59",
      "content": "<p>Thanks for sharing. It would put to use for the next competition</p>",
      "rawMarkdown": "Thanks for sharing. It would put to use for the next competition",
      "votes": null
    },
    {
      "id": "323851",
      "postDate": "05/06/2018 12:10:47",
      "content": "<p>It works everywhere, it is a parameter that doe snot depend on the api you use.  Like all parameters AFAIK.</p>",
      "rawMarkdown": "It works everywhere, it is a parameter that doe snot depend on the api you use.  Like all parameters AFAIK.",
      "votes": null
    },
    {
      "id": "323852",
      "postDate": "05/06/2018 12:12:44",
      "content": "<p>Hi @CPMP, yes I think that this method is practically the same as 'two_stage', but in my case this one consume less RAM in my pipeline. And for use 'two_stage' you have to got already combined dataset on disk, but in this method you can combine a realy very HUGE dataset on disk, and then just loaded necessary parts form it without memory consuming.</p>",
      "rawMarkdown": "Hi @CPMP, yes I think that this method is practically the same as 'two_stage', but in my case this one consume less RAM in my pipeline. And for use 'two_stage' you have to got already combined dataset on disk, but in this method you can combine a realy very HUGE dataset on disk, and then just loaded necessary parts form it without memory consuming.",
      "votes": null
    },
    {
      "id": "323858",
      "postDate": "05/06/2018 12:26:43",
      "content": "<p>@Kruegger,  thanks.  Two stage LB is good enough for now, but you re right, your approach canlead to larger dataset handling.</p>",
      "rawMarkdown": "Kruegger,  thanks.  Two stage LB is good enough for now, but you re right, your approach canlead to larger dataset handling.",
      "votes": null
    },
    {
      "id": "324005",
      "postDate": "05/06/2018 22:25:32",
      "content": "<p>Thanks for sharing! very useful, I'll use it in other competitions.</p>",
      "rawMarkdown": "Thanks for sharing! very useful, I'll use it in other competitions.",
      "votes": null
    },
    {
      "id": "324102",
      "postDate": "05/07/2018 06:55:05",
      "content": "<p>Thanks @Krugger for sharing. It is too late for this one but will surely be using it next time.</p>",
      "rawMarkdown": "Thanks @Krugger for sharing. It is too late for this one but will surely be using it next time.",
      "votes": null
    },
    {
      "id": "324349",
      "postDate": "05/07/2018 16:02:43",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "325590",
      "postDate": "05/08/2018 15:49:39",
      "content": "<p>Bottomline - We have implemented this method in our solution and lift up to 8'th place )\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56325\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56325</a></p>",
      "rawMarkdown": "Bottomline - We have implemented this method in our solution and lift up to 8'th place )\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56325",
      "votes": null
    },
    {
      "id": "3269640",
      "postDate": "08/14/2025 19:07:51",
      "content": "<p>Perfect code to avoid memory leaking!</p>",
      "rawMarkdown": "Perfect code to avoid memory leaking!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3269640,
      "author_name": "xthomasx",
      "author_url": "",
      "post_date": "08/14/2025 19:07:51",
      "content": "<p>Perfect code to avoid memory leaking!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323802,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/06/2018 09:17:03",
      "content": "<p>Hi, thanks for sharing.  Isn't this more or less equivalent to the two_stage of lgb?  Or it is better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 323833,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "05/06/2018 11:11:18",
          "content": "<p>Does the two round parameter works for LGB python package or it’s just for LGB command line?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323851,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/06/2018 12:10:47",
          "content": "<p>It works everywhere, it is a parameter that doe snot depend on the api you use.  Like all parameters AFAIK.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323852,
          "author_name": "kruegger",
          "author_url": "",
          "post_date": "05/06/2018 12:12:44",
          "content": "<p>Hi @CPMP, yes I think that this method is practically the same as 'two_stage', but in my case this one consume less RAM in my pipeline. And for use 'two_stage' you have to got already combined dataset on disk, but in this method you can combine a realy very HUGE dataset on disk, and then just loaded necessary parts form it without memory consuming.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323858,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/06/2018 12:26:43",
          "content": "<p>@Kruegger,  thanks.  Two stage LB is good enough for now, but you re right, your approach canlead to larger dataset handling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 323832,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "05/06/2018 11:10:24",
      "content": "<p>This would be a useful technique. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323838,
      "author_name": "mrbeer",
      "author_url": "",
      "post_date": "05/06/2018 11:20:59",
      "content": "<p>Thanks for sharing. It would put to use for the next competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324005,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/06/2018 22:25:32",
      "content": "<p>Thanks for sharing! very useful, I'll use it in other competitions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324102,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/07/2018 06:55:05",
      "content": "<p>Thanks @Krugger for sharing. It is too late for this one but will surely be using it next time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324349,
      "author_name": "sawseen",
      "author_url": "",
      "post_date": "05/07/2018 16:02:43",
      "content": "<p>Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325590,
      "author_name": "kruegger",
      "author_url": "",
      "post_date": "05/08/2018 15:49:39",
      "content": "<p>Bottomline - We have implemented this method in our solution and lift up to 8'th place )\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56325\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56325</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "323767": "This competition is very hard according to large data, so all kaggler's try to use some tricks to train with more data in memory constraint environment. Thanks to all who shared their approaches - it helps a lot! Here I want to add one more trick to bag of community solution - how to combine several files with group of pickled attrs to one train/val without memory spikes.  \n\nLet we have separate group of attrs pickled in different files 'attrs_?????' in directory './data/'. The each file is concatenated [train,test] part of train dataframe.\n\n    split_train = 11111111     # border between train/test in pickle (length of our train)\n    start_skip = 0   # if we want strip our train\n    n_attrs = 25  # as is\n    \n    pickle_list = ['attrs_xxxx',   'attrs_base_cnt',   'attrs_nunique']   # list of pickled attrs\n    del_cols = [  'click_time',  'day', ]  # attrs to be deleted in final train\n\nAt first we create memmap file where we will store our combined train dataframe\n\n    si = 0   \n    mmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='w+', shape=(split_train, n_attrs))\n\ndefine helper function\n\n    def load_attrs(fname, data_dir='./data/'):\n        fname = data_dir + fname + '.pkl'\n        print('loading {}... '.format(fname))\n        return pd.read_pickle(fname)\n\nand convert our parts (of attrs) to appropriate position in the memmap file. Note - we have to save columns of our combined dataframe as memmap is just numpy array (less or more :) ) and it doesn't have information about columns.\n\n    columns = []\n    for pkl in pickle_list:  \n        _temp = load_attrs(pkl)\n    \n        _columns  = [x for x in _temp.columns if x not in del_cols]  \n        columns = columns + _columns\n    \n        nodel_ind = [_temp.columns.tolist().index(x) for x in _temp.columns if x not in del_cols]  \n        \n        _temp = _temp.iloc[start_skip:split_train, nodel_ind]  \n        \n        ei = _temp.values.shape[1]  \n        mmap[:, si:si+ei] = _temp.values  \n        si += ei  \n         \n        del _temp  \n        gc.collect()\n\nthen we flushed our memmap file and close it.\n\n    mmap.flush()\n    del mmap\n    gc.collect()\n\n\nreopen our memmap in 'read-only' mode and extract train/val as numpy array.\n\n    mmap = np.memmap(r'./data/mmap_train.mmp', dtype='float32', mode='r', shape=(split_train, n_attrs))\n    _train = np.array(mmap[start_skip:-val_size])\n    _val = np.array(mmap[-val_size:])\n\nget y\n\n    _train = _train[:, columns.index('is_attributed')]\n    _val = _val[:, columns.index('is_attributed')]\n\nprepare set of columns to use in training and get indexes of them.\n\n    use_columns = [ 'app',  'channel',  'device',  'os',  'hour', 'ip_cnt']\n    muse_columns = [columns.index(x) for x in use_columns]\n\nprepare lgb.Datasets\n\n    d_train = _train[:, muse_columns]\n    xgtrain = lgb.Dataset(d_train, label=yy_train, **dataset_params)\n    d_val = mXX_val[:, muse_columns]\n    xgvalid = lgb.Dataset(d_val, label=yy_val, reference=xgtrain, **dataset_params)\n\ntrain\n\n\n    bst = lgb.train(model_params, xgtrain, valid_sets=[xgvalid], valid_names=['valid'], evals_result=evals_results, **fit_params)\n\nand voila! No spikes, no memory leaking, no ARGHHHH )  \nHope this hepls.\n\nHappy kaggling!  \nKruegger\n\nP.S. you can upvote if you find this useful )",
    "323802": "Hi, thanks for sharing.  Isn't this more or less equivalent to the two_stage of lgb?  Or it is better?",
    "323832": "This would be a useful technique.",
    "323833": "Does the two round parameter works for LGB python package or it’s just for LGB command line?",
    "323838": "Thanks for sharing. It would put to use for the next competition",
    "323851": "It works everywhere, it is a parameter that doe snot depend on the api you use.  Like all parameters AFAIK.",
    "323852": "Hi @CPMP, yes I think that this method is practically the same as 'two_stage', but in my case this one consume less RAM in my pipeline. And for use 'two_stage' you have to got already combined dataset on disk, but in this method you can combine a realy very HUGE dataset on disk, and then just loaded necessary parts form it without memory consuming.",
    "323858": "Kruegger,  thanks.  Two stage LB is good enough for now, but you re right, your approach canlead to larger dataset handling.",
    "324005": "Thanks for sharing! very useful, I'll use it in other competitions.",
    "324102": "Thanks @Krugger for sharing. It is too late for this one but will surely be using it next time.",
    "324349": "Thanks a lot!",
    "325590": "Bottomline - We have implemented this method in our solution and lift up to 8'th place )\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56325",
    "3269640": "Perfect code to avoid memory leaking!"
  },
  "source": "meta"
}