{
  "id": 55030,
  "title": "28 New Features - 0.9803 LB Score [Updated - v5]",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55030",
  "author_name": "Samrat Pandiri",
  "post_date": "2018-04-21T06:03:53.076000",
  "votes": 18,
  "comment_count": 40,
  "views": 0,
  "content": "<p><strong>Update - May 04th</strong> </p>\n\n<ul>\n<li>Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.</li>\n</ul>\n\n<p>Giving one last try..... adding 2 new features and removed 3 features</p>\n\n<p><strong>Update - April 26th</strong> </p>\n\n<ul>\n<li>Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.</li>\n</ul>\n\n<p><strong>Update - April 25th</strong></p>\n\n<ul>\n<li>Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.</li>\n<li>Looking to work on the feature importance to remove some under performing features and add some new features.</li>\n</ul>\n\n<p><strong>Update - April 23rd</strong></p>\n\n<p>Added 4 more features and the local validation jumped to 0.990979. The LB score also improved to 0.9800.</p>\n\n<p>Apart from the base features, added 9 new features to get a LB score of 0.9798.</p>\n\n<p>train size:  182403889\nvalid size:  2500000 (tail rows)</p>\n\n<p>Training until validation scores don't improve for 50 rounds.\n[10]    train's auc: 0.971521   valid's auc: 0.978039\n[20]    train's auc: 0.977205   valid's auc: 0.981515\n[30]    train's auc: 0.979879   valid's auc: 0.984628\n[40]    train's auc: 0.981305   valid's auc: 0.987058\n[50]    train's auc: 0.982025   valid's auc: 0.987659\n[60]    train's auc: 0.982491   valid's auc: 0.988301\n[70]    train's auc: 0.982942   valid's auc: 0.988885\n[80]    train's auc: 0.983197   valid's auc: 0.989289\n[90]    train's auc: 0.983447   valid's auc: 0.989419\n[100]   train's auc: 0.983683   valid's auc: 0.989379\n[110]   train's auc: 0.98386    valid's auc: 0.989594\n[120]   train's auc: 0.983984   valid's auc: 0.989651\n[130]   train's auc: 0.984101   valid's auc: 0.989784\n[140]   train's auc: 0.984219   valid's auc: 0.989845\n[150]   train's auc: 0.984317   valid's auc: 0.989851\n[160]   train's auc: 0.984398   valid's auc: 0.990049\n[170]   train's auc: 0.984523   valid's auc: 0.99027\n[180]   train's auc: 0.984617   valid's auc: 0.990348\n[190]   train's auc: 0.984713   valid's auc: 0.990424\n[200]   train's auc: 0.984799   valid's auc: 0.990447\n[210]   train's auc: 0.984853   valid's auc: 0.990475\n[220]   train's auc: 0.984915   valid's auc: 0.990588\n[230]   train's auc: 0.984979   valid's auc: 0.990649\n[240]   train's auc: 0.985035   valid's auc: 0.990668\n[250]   train's auc: 0.985089   valid's auc: 0.990664\n[260]   train's auc: 0.985126   valid's auc: 0.990703\n[270]   train's auc: 0.985166   valid's auc: 0.990704\n[280]   train's auc: 0.985209   valid's auc: 0.990698\n[290]   train's auc: 0.985251   valid's auc: 0.990775\n[300]   train's auc: 0.985285   valid's auc: 0.990797\n[310]   train's auc: 0.985315   valid's auc: 0.99079\n[320]   train's auc: 0.985351   valid's auc: 0.990758\n[330]   train's auc: 0.985379   valid's auc: 0.990737\n[340]   train's auc: 0.985418   valid's auc: 0.990764\n[350]   train's auc: 0.98545    valid's auc: 0.99076\nEarly stopping, best iteration is:\n[307]   train's auc: 0.985306   valid's auc: 0.990801</p>\n\n<p>Using 32GB Machine with 100GB Swap Space on 2 SSD's. I usually starts the script at night and wake up in the morning to see the beautiful results :)</p>\n\n<p>PS: 7 Features gave me a LB score of 0.9788.</p>\n\n<p>The next step is to add 3 more new features and see if it can get a score of 0.98XX</p>",
  "messages": [
    {
      "id": 317293,
      "postDate": "2018-04-21T06:03:53.077Z",
      "content": "<p><strong>Update - May 04th</strong> </p>\n\n<ul>\n<li>Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.</li>\n</ul>\n\n<p>Giving one last try..... adding 2 new features and removed 3 features</p>\n\n<p><strong>Update - April 26th</strong> </p>\n\n<ul>\n<li>Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.</li>\n</ul>\n\n<p><strong>Update - April 25th</strong></p>\n\n<ul>\n<li>Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.</li>\n<li>Looking to work on the feature importance to remove some under performing features and add some new features.</li>\n</ul>\n\n<p><strong>Update - April 23rd</strong></p>\n\n<p>Added 4 more features and the local validation jumped to 0.990979. The LB score also improved to 0.9800.</p>\n\n<p>Apart from the base features, added 9 new features to get a LB score of 0.9798.</p>\n\n<p>train size:  182403889\nvalid size:  2500000 (tail rows)</p>\n\n<p>Training until validation scores don't improve for 50 rounds.\n[10]    train's auc: 0.971521   valid's auc: 0.978039\n[20]    train's auc: 0.977205   valid's auc: 0.981515\n[30]    train's auc: 0.979879   valid's auc: 0.984628\n[40]    train's auc: 0.981305   valid's auc: 0.987058\n[50]    train's auc: 0.982025   valid's auc: 0.987659\n[60]    train's auc: 0.982491   valid's auc: 0.988301\n[70]    train's auc: 0.982942   valid's auc: 0.988885\n[80]    train's auc: 0.983197   valid's auc: 0.989289\n[90]    train's auc: 0.983447   valid's auc: 0.989419\n[100]   train's auc: 0.983683   valid's auc: 0.989379\n[110]   train's auc: 0.98386    valid's auc: 0.989594\n[120]   train's auc: 0.983984   valid's auc: 0.989651\n[130]   train's auc: 0.984101   valid's auc: 0.989784\n[140]   train's auc: 0.984219   valid's auc: 0.989845\n[150]   train's auc: 0.984317   valid's auc: 0.989851\n[160]   train's auc: 0.984398   valid's auc: 0.990049\n[170]   train's auc: 0.984523   valid's auc: 0.99027\n[180]   train's auc: 0.984617   valid's auc: 0.990348\n[190]   train's auc: 0.984713   valid's auc: 0.990424\n[200]   train's auc: 0.984799   valid's auc: 0.990447\n[210]   train's auc: 0.984853   valid's auc: 0.990475\n[220]   train's auc: 0.984915   valid's auc: 0.990588\n[230]   train's auc: 0.984979   valid's auc: 0.990649\n[240]   train's auc: 0.985035   valid's auc: 0.990668\n[250]   train's auc: 0.985089   valid's auc: 0.990664\n[260]   train's auc: 0.985126   valid's auc: 0.990703\n[270]   train's auc: 0.985166   valid's auc: 0.990704\n[280]   train's auc: 0.985209   valid's auc: 0.990698\n[290]   train's auc: 0.985251   valid's auc: 0.990775\n[300]   train's auc: 0.985285   valid's auc: 0.990797\n[310]   train's auc: 0.985315   valid's auc: 0.99079\n[320]   train's auc: 0.985351   valid's auc: 0.990758\n[330]   train's auc: 0.985379   valid's auc: 0.990737\n[340]   train's auc: 0.985418   valid's auc: 0.990764\n[350]   train's auc: 0.98545    valid's auc: 0.99076\nEarly stopping, best iteration is:\n[307]   train's auc: 0.985306   valid's auc: 0.990801</p>\n\n<p>Using 32GB Machine with 100GB Swap Space on 2 SSD's. I usually starts the script at night and wake up in the morning to see the beautiful results :)</p>\n\n<p>PS: 7 Features gave me a LB score of 0.9788.</p>\n\n<p>The next step is to add 3 more new features and see if it can get a score of 0.98XX</p>",
      "rawMarkdown": "**Update - May 04th** \n\n- Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.\n\nGiving one last try..... adding 2 new features and removed 3 features\n\n**Update - April 26th** \n\n- Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.\n\n**Update - April 25th**\n\n- Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.\n- Looking to work on the feature importance to remove some under performing features and add some new features.\n\n**Update - April 23rd**\n\nAdded 4 more features and the local validation jumped to 0.990979. The LB score also improved to 0.9800.\n\nApart from the base features, added 9 new features to get a LB score of 0.9798.\n\ntrain size:  182403889\nvalid size:  2500000 (tail rows)\n\nTraining until validation scores don't improve for 50 rounds.\n[10]    train's auc: 0.971521   valid's auc: 0.978039\n[20]    train's auc: 0.977205   valid's auc: 0.981515\n[30]    train's auc: 0.979879   valid's auc: 0.984628\n[40]    train's auc: 0.981305   valid's auc: 0.987058\n[50]    train's auc: 0.982025   valid's auc: 0.987659\n[60]    train's auc: 0.982491   valid's auc: 0.988301\n[70]    train's auc: 0.982942   valid's auc: 0.988885\n[80]    train's auc: 0.983197   valid's auc: 0.989289\n[90]    train's auc: 0.983447   valid's auc: 0.989419\n[100]   train's auc: 0.983683   valid's auc: 0.989379\n[110]   train's auc: 0.98386    valid's auc: 0.989594\n[120]   train's auc: 0.983984   valid's auc: 0.989651\n[130]   train's auc: 0.984101   valid's auc: 0.989784\n[140]   train's auc: 0.984219   valid's auc: 0.989845\n[150]   train's auc: 0.984317   valid's auc: 0.989851\n[160]   train's auc: 0.984398   valid's auc: 0.990049\n[170]   train's auc: 0.984523   valid's auc: 0.99027\n[180]   train's auc: 0.984617   valid's auc: 0.990348\n[190]   train's auc: 0.984713   valid's auc: 0.990424\n[200]   train's auc: 0.984799   valid's auc: 0.990447\n[210]   train's auc: 0.984853   valid's auc: 0.990475\n[220]   train's auc: 0.984915   valid's auc: 0.990588\n[230]   train's auc: 0.984979   valid's auc: 0.990649\n[240]   train's auc: 0.985035   valid's auc: 0.990668\n[250]   train's auc: 0.985089   valid's auc: 0.990664\n[260]   train's auc: 0.985126   valid's auc: 0.990703\n[270]   train's auc: 0.985166   valid's auc: 0.990704\n[280]   train's auc: 0.985209   valid's auc: 0.990698\n[290]   train's auc: 0.985251   valid's auc: 0.990775\n[300]   train's auc: 0.985285   valid's auc: 0.990797\n[310]   train's auc: 0.985315   valid's auc: 0.99079\n[320]   train's auc: 0.985351   valid's auc: 0.990758\n[330]   train's auc: 0.985379   valid's auc: 0.990737\n[340]   train's auc: 0.985418   valid's auc: 0.990764\n[350]   train's auc: 0.98545    valid's auc: 0.99076\nEarly stopping, best iteration is:\n[307]   train's auc: 0.985306   valid's auc: 0.990801\n\nUsing 32GB Machine with 100GB Swap Space on 2 SSD's. I usually starts the script at night and wake up in the morning to see the beautiful results :)\n\nPS: 7 Features gave me a LB score of 0.9788.\n\nThe next step is to add 3 more new features and see if it can get a score of 0.98XX",
      "votes": 18
    },
    {
      "id": 317674,
      "postDate": "2018-04-22T07:50:15.087Z",
      "content": "<p>@Samrat\nThanks for the share.\n50 rounds?\nI use 20, do you find cases where after 20 rounds of no better AUC suddenly it jumps up again?</p>\n\n<ul>\n<li>It's amusing that your Validation AUC is higher than the Training AUC, Don't you prefer to take a specific planned set of data instead of just the tail? Like day9 \\ day8 \\ specific hours?</li>\n</ul>",
      "rawMarkdown": "@Samrat\nThanks for the share.\n50 rounds?\nI use 20, do you find cases where after 20 rounds of no better AUC suddenly it jumps up again?\n\n+ It's amusing that your Validation AUC is higher than the Training AUC, Don't you prefer to take a specific planned set of data instead of just the tail? Like day9 \\ day8 \\ specific hours?",
      "votes": 1,
      "replies": [
        {
          "id": 317736,
          "postDate": "2018-04-22T12:22:29.547Z",
          "content": "<p>Thanks Amir for your inputs. This is my first competition and I'm trying out many new things and also learning a lot from the community..\nI started with using the tail data for the validation and I'm seeing a positive correlation with the local validation amd public lb score... so, I'm continuing with the same... Will try using the last day's data for validation.</p>",
          "rawMarkdown": "Thanks Amir for your inputs. This is my first competition and I'm trying out many new things and also learning a lot from the community..\nI started with using the tail data for the validation and I'm seeing a positive correlation with the local validation amd public lb score... so, I'm continuing with the same... Will try using the last day's data for validation.",
          "votes": 1
        },
        {
          "id": 317752,
          "postDate": "2018-04-22T13:18:22.300Z",
          "content": "<p>Hi @Samrat,\nIt's also my first competition,\nI also played around with quite a few different CV methods.\nMy latest method, after following some discussions about it, is ignoring the closeness of CV AUC and Public LB and just set a good stable CV set, it did improve my LB score.\nAlthough it hurts to have such high difference of AUC between the 2, it's supposed to generalize better.\nThis competition has some really odd training + test sets that pushes Kagglers to overfit to a specific set of data that might not generalize well in Private LB or the actual data outside this competition.</p>\n\n<p>Let me know how it goes with that new validation set =)</p>",
          "rawMarkdown": "Hi @Samrat,\nIt's also my first competition,\nI also played around with quite a few different CV methods.\nMy latest method, after following some discussions about it, is ignoring the closeness of CV AUC and Public LB and just set a good stable CV set, it did improve my LB score.\nAlthough it hurts to have such high difference of AUC between the 2, it's supposed to generalize better.\nThis competition has some really odd training + test sets that pushes Kagglers to overfit to a specific set of data that might not generalize well in Private LB or the actual data outside this competition.\n\nLet me know how it goes with that new validation set =)",
          "votes": 1
        },
        {
          "id": 318113,
          "postDate": "2018-04-23T07:48:36.307Z",
          "content": "<p>Sure Amir.. let me give that a try.. </p>",
          "rawMarkdown": "Sure Amir.. let me give that a try.. "
        }
      ]
    },
    {
      "id": 319979,
      "postDate": "2018-04-27T08:21:17.173Z",
      "content": "<p>I think that you have very small valid set, and delta between valid score and LB score is really very big...</p>",
      "rawMarkdown": "I think that you have very small valid set, and delta between valid score and LB score is really very big...",
      "votes": 2,
      "replies": [
        {
          "id": 320045,
          "postDate": "2018-04-27T10:52:09.010Z",
          "content": "<p>If using the last 5M samples of the train data, that what happens. There is a strong relation of AUC as a function of time of day. Coincidentally, the last 5M samples are in a very good time.\nThat's why I moved to using day 9 as evaluation.</p>",
          "rawMarkdown": "If using the last 5M samples of the train data, that what happens. There is a strong relation of AUC as a function of time of day. Coincidentally, the last 5M samples are in a very good time.\nThat's why I moved to using day 9 as evaluation.",
          "votes": 3
        },
        {
          "id": 321339,
          "postDate": "2018-05-01T04:20:23.620Z",
          "content": "<p>@Yair Totally agreed! </p>",
          "rawMarkdown": "@Yair Totally agreed! "
        }
      ]
    },
    {
      "id": 323006,
      "postDate": "2018-05-04T06:00:48.680Z",
      "content": "<p>Ensemble with another model of mine gives LB score of 0.9806.</p>\n\n<p>That's impressive. </p>",
      "rawMarkdown": "Ensemble with another model of mine gives LB score of 0.9806.\n\nThat's impressive. ",
      "replies": [
        {
          "id": 323011,
          "postDate": "2018-05-04T06:18:27.667Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 323066,
          "postDate": "2018-05-04T09:16:17.920Z",
          "content": "<p>For now, I only use single lgb model. Having no idea about how to ensemble.</p>",
          "rawMarkdown": "For now, I only use single lgb model. Having no idea about how to ensemble."
        },
        {
          "id": 323083,
          "postDate": "2018-05-04T10:23:58.090Z",
          "content": "<p>@Liu Jilong\nThis is the time to read about it and try!\nGoing to spend the rest of the competition on learning how to ensemble and see how it goes.</p>",
          "rawMarkdown": "@Liu Jilong\nThis is the time to read about it and try!\nGoing to spend the rest of the competition on learning how to ensemble and see how it goes."
        }
      ]
    },
    {
      "id": 322974,
      "postDate": "2018-05-04T03:58:52.677Z",
      "content": "<p><strong>Update - May 04th</strong> </p>\n\n<ul>\n<li>Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.</li>\n</ul>\n\n<p>Giving one last try..... adding 2 new features and removed 3 features</p>",
      "rawMarkdown": "**Update - May 04th** \n\n- Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.\n\nGiving one last try..... adding 2 new features and removed 3 features",
      "replies": [
        {
          "id": 322978,
          "postDate": "2018-05-04T04:17:27.443Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 322172,
      "postDate": "2018-05-02T13:59:28.053Z",
      "content": "<p>hi, could you tell me how did you generate your local cv dataset?</p>",
      "rawMarkdown": "hi, could you tell me how did you generate your local cv dataset?",
      "replies": [
        {
          "id": 322178,
          "postDate": "2018-05-02T14:09:58.500Z",
          "content": "<p>I'm just using the last 2.5M rows.... Not a good strategy thou...</p>",
          "rawMarkdown": "I'm just using the last 2.5M rows.... Not a good strategy thou..."
        },
        {
          "id": 322192,
          "postDate": "2018-05-02T14:29:04.397Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 322462,
          "postDate": "2018-05-03T02:07:50.137Z",
          "content": "<p>train dataset is too huge for me....</p>",
          "rawMarkdown": "train dataset is too huge for me...."
        }
      ]
    },
    {
      "id": 319886,
      "postDate": "2018-04-27T03:53:24.313Z",
      "content": "<p>hello,thanks for your sharing.How much memory did you use? Did you read csv in chunk or read them all ? When I read them all together, I got a memoryerr exception,but    the total memory is 120GB.I am really comfused.</p>",
      "rawMarkdown": "hello,thanks for your sharing.How much memory did you use? Did you read csv in chunk or read them all ? When I read them all together, I got a memoryerr exception,but    the total memory is 120GB.I am really comfused."
    },
    {
      "id": 319715,
      "postDate": "2018-04-26T17:15:19.767Z",
      "content": "<p><strong>Update</strong> - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.</p>",
      "rawMarkdown": "**Update** - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.",
      "replies": [
        {
          "id": 319853,
          "postDate": "2018-04-27T00:51:36.420Z",
          "content": "<p>could you tell me what kind of features you add?</p>",
          "rawMarkdown": "could you tell me what kind of features you add?"
        },
        {
          "id": 319860,
          "postDate": "2018-04-27T01:46:29.880Z",
          "content": "<p>@MengYe Current I have couple of time delta features and others are aggregate features...</p>",
          "rawMarkdown": "@MengYe Current I have couple of time delta features and others are aggregate features..."
        },
        {
          "id": 319863,
          "postDate": "2018-04-27T01:51:55.657Z",
          "content": "<p>are confRate features having high importance in your model, target encoding is utilized right?</p>",
          "rawMarkdown": "are confRate features having high importance in your model, target encoding is utilized right?"
        }
      ]
    },
    {
      "id": 319232,
      "postDate": "2018-04-25T14:58:41.607Z",
      "content": "<p><strong>Update</strong> - Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.</p>",
      "rawMarkdown": "**Update** - Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.",
      "replies": [
        {
          "id": 319648,
          "postDate": "2018-04-26T14:23:04.457Z",
          "content": "<p>are you experimenting with count features or target encoding features or time delta features, what is working best for you?\nAlso , has anyone started with hyperparameter optimization yet?</p>",
          "rawMarkdown": "are you experimenting with count features or target encoding features or time delta features, what is working best for you?\nAlso , has anyone started with hyperparameter optimization yet?"
        },
        {
          "id": 319686,
          "postDate": "2018-04-26T15:40:02.487Z",
          "content": "<p>Currently I'm working on count features only... after removing some under performing features will work on time delta features..</p>",
          "rawMarkdown": "Currently I'm working on count features only... after removing some under performing features will work on time delta features.."
        }
      ]
    },
    {
      "id": 318537,
      "postDate": "2018-04-24T02:47:36.837Z",
      "content": "<p>9 new features, 0.9788</p>",
      "rawMarkdown": "9 new features, 0.9788"
    },
    {
      "id": 318464,
      "postDate": "2018-04-23T22:04:47.753Z",
      "content": "<p>Thanks for the share, I am wondering if &lt; 0.98XX is the limit of single model and you just give the answer.</p>",
      "rawMarkdown": "Thanks for the share, I am wondering if &lt; 0.98XX is the limit of single model and you just give the answer.",
      "replies": [
        {
          "id": 318528,
          "postDate": "2018-04-24T02:13:54.923Z",
          "content": "<p>I dont think so.... I see people having better scores with a single model.. I added few new features and removed some in my current code.. Will update you if I see better results with a single model...</p>",
          "rawMarkdown": "I dont think so.... I see people having better scores with a single model.. I added few new features and removed some in my current code.. Will update you if I see better results with a single model..."
        }
      ]
    },
    {
      "id": 318010,
      "postDate": "2018-04-23T03:17:15.187Z",
      "content": "<p>and this my first competitions too :)</p>",
      "rawMarkdown": "and this my first competitions too :)"
    },
    {
      "id": 318009,
      "postDate": "2018-04-23T03:16:43.950Z",
      "content": "<p>I use 100 million data, got 0.9791, 9 Features ; I think use more data get higher score。 but I dont have more memory in my laptop</p>",
      "rawMarkdown": "I use 100 million data, got 0.9791, 9 Features ; I think use more data get higher score。 but I dont have more memory in my laptop",
      "replies": [
        {
          "id": 318022,
          "postDate": "2018-04-23T03:47:46.677Z",
          "content": "<p>@Johnny Did you try using swap space so that you can easily push in extra 10-20 M rows....\nI have 2 SSD's and using the swap space to push in more rows...</p>",
          "rawMarkdown": "@Johnny Did you try using swap space so that you can easily push in extra 10-20 M rows....\nI have 2 SSD's and using the swap space to push in more rows...",
          "votes": 1
        },
        {
          "id": 318520,
          "postDate": "2018-04-24T01:35:54.037Z",
          "content": "<p>Yes, I use 130GB swap sapce</p>",
          "rawMarkdown": "Yes, I use 130GB swap sapce"
        },
        {
          "id": 318522,
          "postDate": "2018-04-24T01:40:45.573Z",
          "content": "<p>I use mac air laptop, 8GB Ram. :)</p>",
          "rawMarkdown": "I use mac air laptop, 8GB Ram. :)",
          "votes": 1
        },
        {
          "id": 318529,
          "postDate": "2018-04-24T02:18:26.240Z",
          "content": "<p>Probably you can try Azure or GCP to see if your model does better with complete rows... I think the free credits should be good enough for the competition...</p>",
          "rawMarkdown": "Probably you can try Azure or GCP to see if your model does better with complete rows... I think the free credits should be good enough for the competition...",
          "votes": 1
        },
        {
          "id": 318550,
          "postDate": "2018-04-24T03:29:05.350Z",
          "content": "<p>Thanks so much，I will try recently ：）</p>",
          "rawMarkdown": "Thanks so much，I will try recently ：）"
        },
        {
          "id": 318587,
          "postDate": "2018-04-24T05:23:39.803Z",
          "content": "<p>Hey, is there any tutorial about how to use swap space in Mac? I also have 8GB RAM and kernel will dead when I use lines more than 75M. Thanks =]</p>",
          "rawMarkdown": "Hey, is there any tutorial about how to use swap space in Mac? I also have 8GB RAM and kernel will dead when I use lines more than 75M. Thanks =]"
        },
        {
          "id": 319867,
          "postDate": "2018-04-27T02:02:15.630Z",
          "content": "<p>I don't use any setting, mac OS will use swap space automatically.As long as u have enough free space, I deleted a lot of files to have 130g space in my SSD. </p>",
          "rawMarkdown": "I don't use any setting, mac OS will use swap space automatically.As long as u have enough free space, I deleted a lot of files to have 130g space in my SSD. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 317615,
      "postDate": "2018-04-22T01:54:16.477Z",
      "content": "<p>PS: 7 Features gave me a LB score of 0.9788.</p>\n\n<p>Is the model using only 7 features to get 0.9788 ?\nThanks .</p>",
      "rawMarkdown": "PS: 7 Features gave me a LB score of 0.9788.\n\nIs the model using only 7 features to get 0.9788 ?\nThanks .",
      "replies": [
        {
          "id": 317628,
          "postDate": "2018-04-22T03:19:52.050Z",
          "content": "<p>I mean new features..</p>",
          "rawMarkdown": "I mean new features..",
          "votes": 1
        }
      ]
    },
    {
      "id": 317294,
      "postDate": "2018-04-21T06:11:10.673Z",
      "content": "<p>Thanks for the share.</p>",
      "rawMarkdown": "Thanks for the share."
    }
  ],
  "comments": [
    {
      "id": 317674,
      "author_name": "AmirH",
      "author_url": "",
      "post_date": "2018-04-22T07:50:15.087000",
      "content": "<p>@Samrat\nThanks for the share.\n50 rounds?\nI use 20, do you find cases where after 20 rounds of no better AUC suddenly it jumps up again?</p>\n\n<ul>\n<li>It's amusing that your Validation AUC is higher than the Training AUC, Don't you prefer to take a specific planned set of data instead of just the tail? Like day9 \\ day8 \\ specific hours?</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 317736,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-22T12:22:29.547000",
          "content": "<p>Thanks Amir for your inputs. This is my first competition and I'm trying out many new things and also learning a lot from the community..\nI started with using the tail data for the validation and I'm seeing a positive correlation with the local validation amd public lb score... so, I'm continuing with the same... Will try using the last day's data for validation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 317752,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-22T13:18:22.300000",
          "content": "<p>Hi @Samrat,\nIt's also my first competition,\nI also played around with quite a few different CV methods.\nMy latest method, after following some discussions about it, is ignoring the closeness of CV AUC and Public LB and just set a good stable CV set, it did improve my LB score.\nAlthough it hurts to have such high difference of AUC between the 2, it's supposed to generalize better.\nThis competition has some really odd training + test sets that pushes Kagglers to overfit to a specific set of data that might not generalize well in Private LB or the actual data outside this competition.</p>\n\n<p>Let me know how it goes with that new validation set =)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318113,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-23T07:48:36.307000",
          "content": "<p>Sure Amir.. let me give that a try.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319979,
      "author_name": "Nikita Varganov",
      "author_url": "",
      "post_date": "2018-04-27T08:21:17.173000",
      "content": "<p>I think that you have very small valid set, and delta between valid score and LB score is really very big...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 320045,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-04-27T10:52:09.010000",
          "content": "<p>If using the last 5M samples of the train data, that what happens. There is a strong relation of AUC as a function of time of day. Coincidentally, the last 5M samples are in a very good time.\nThat's why I moved to using day 9 as evaluation.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 321339,
          "author_name": "Wenjie Bai",
          "author_url": "",
          "post_date": "2018-05-01T04:20:23.620000",
          "content": "<p>@Yair Totally agreed! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 323006,
      "author_name": "Liu Jilong",
      "author_url": "",
      "post_date": "2018-05-04T06:00:48.680000",
      "content": "<p>Ensemble with another model of mine gives LB score of 0.9806.</p>\n\n<p>That's impressive. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 323011,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-04T06:18:27.667000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323066,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-05-04T09:16:17.920000",
          "content": "<p>For now, I only use single lgb model. Having no idea about how to ensemble.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323083,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-05-04T10:23:58.090000",
          "content": "<p>@Liu Jilong\nThis is the time to read about it and try!\nGoing to spend the rest of the competition on learning how to ensemble and see how it goes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 322974,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-05-04T03:58:52.677000",
      "content": "<p><strong>Update - May 04th</strong> </p>\n\n<ul>\n<li>Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.</li>\n</ul>\n\n<p>Giving one last try..... adding 2 new features and removed 3 features</p>",
      "votes": 0,
      "replies": [
        {
          "id": 322978,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-04T04:17:27.443000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 322172,
      "author_name": "Marvin Free",
      "author_url": "",
      "post_date": "2018-05-02T13:59:28.053000",
      "content": "<p>hi, could you tell me how did you generate your local cv dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 322178,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-05-02T14:09:58.500000",
          "content": "<p>I'm just using the last 2.5M rows.... Not a good strategy thou...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322192,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-02T14:29:04.397000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322462,
          "author_name": "Marvin Free",
          "author_url": "",
          "post_date": "2018-05-03T02:07:50.137000",
          "content": "<p>train dataset is too huge for me....</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319886,
      "author_name": "twoone",
      "author_url": "",
      "post_date": "2018-04-27T03:53:24.313000",
      "content": "<p>hello,thanks for your sharing.How much memory did you use? Did you read csv in chunk or read them all ? When I read them all together, I got a memoryerr exception,but    the total memory is 120GB.I am really comfused.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319715,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-04-26T17:15:19.767000",
      "content": "<p><strong>Update</strong> - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 319853,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-04-27T00:51:36.420000",
          "content": "<p>could you tell me what kind of features you add?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319860,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-27T01:46:29.880000",
          "content": "<p>@MengYe Current I have couple of time delta features and others are aggregate features...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319863,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-27T01:51:55.657000",
          "content": "<p>are confRate features having high importance in your model, target encoding is utilized right?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319232,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-04-25T14:58:41.607000",
      "content": "<p><strong>Update</strong> - Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 319648,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-26T14:23:04.457000",
          "content": "<p>are you experimenting with count features or target encoding features or time delta features, what is working best for you?\nAlso , has anyone started with hyperparameter optimization yet?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319686,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-26T15:40:02.487000",
          "content": "<p>Currently I'm working on count features only... after removing some under performing features will work on time delta features..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 318537,
      "author_name": "YulinGUO",
      "author_url": "",
      "post_date": "2018-04-24T02:47:36.837000",
      "content": "<p>9 new features, 0.9788</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 318464,
      "author_name": "Toaru",
      "author_url": "",
      "post_date": "2018-04-23T22:04:47.753000",
      "content": "<p>Thanks for the share, I am wondering if &lt; 0.98XX is the limit of single model and you just give the answer.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 318528,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-24T02:13:54.923000",
          "content": "<p>I dont think so.... I see people having better scores with a single model.. I added few new features and removed some in my current code.. Will update you if I see better results with a single model...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 318010,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-04-23T03:17:15.187000",
      "content": "<p>and this my first competitions too :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 318009,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-04-23T03:16:43.950000",
      "content": "<p>I use 100 million data, got 0.9791, 9 Features ; I think use more data get higher score。 but I dont have more memory in my laptop</p>",
      "votes": 0,
      "replies": [
        {
          "id": 318022,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-23T03:47:46.677000",
          "content": "<p>@Johnny Did you try using swap space so that you can easily push in extra 10-20 M rows....\nI have 2 SSD's and using the swap space to push in more rows...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318520,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-24T01:35:54.037000",
          "content": "<p>Yes, I use 130GB swap sapce</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318522,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-24T01:40:45.573000",
          "content": "<p>I use mac air laptop, 8GB Ram. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318529,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-24T02:18:26.240000",
          "content": "<p>Probably you can try Azure or GCP to see if your model does better with complete rows... I think the free credits should be good enough for the competition...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318550,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-24T03:29:05.350000",
          "content": "<p>Thanks so much，I will try recently ：）</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318587,
          "author_name": "Perfect Is Shit",
          "author_url": "",
          "post_date": "2018-04-24T05:23:39.803000",
          "content": "<p>Hey, is there any tutorial about how to use swap space in Mac? I also have 8GB RAM and kernel will dead when I use lines more than 75M. Thanks =]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319867,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-27T02:02:15.630000",
          "content": "<p>I don't use any setting, mac OS will use swap space automatically.As long as u have enough free space, I deleted a lot of files to have 130g space in my SSD. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 317615,
      "author_name": "Nomoreaday",
      "author_url": "",
      "post_date": "2018-04-22T01:54:16.477000",
      "content": "<p>PS: 7 Features gave me a LB score of 0.9788.</p>\n\n<p>Is the model using only 7 features to get 0.9788 ?\nThanks .</p>",
      "votes": 0,
      "replies": [
        {
          "id": 317628,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-22T03:19:52.050000",
          "content": "<p>I mean new features..</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 317294,
      "author_name": "Pruthvi Priya",
      "author_url": "",
      "post_date": "2018-04-21T06:11:10.673000",
      "content": "<p>Thanks for the share.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "317293": "**Update - May 04th** \n\n- Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.\n\nGiving one last try..... adding 2 new features and removed 3 features\n\n**Update - April 26th** \n\n- Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.\n\n**Update - April 25th**\n\n- Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.\n- Looking to work on the feature importance to remove some under performing features and add some new features.\n\n**Update - April 23rd**\n\nAdded 4 more features and the local validation jumped to 0.990979. The LB score also improved to 0.9800.\n\nApart from the base features, added 9 new features to get a LB score of 0.9798.\n\ntrain size:  182403889\nvalid size:  2500000 (tail rows)\n\nTraining until validation scores don't improve for 50 rounds.\n[10]    train's auc: 0.971521   valid's auc: 0.978039\n[20]    train's auc: 0.977205   valid's auc: 0.981515\n[30]    train's auc: 0.979879   valid's auc: 0.984628\n[40]    train's auc: 0.981305   valid's auc: 0.987058\n[50]    train's auc: 0.982025   valid's auc: 0.987659\n[60]    train's auc: 0.982491   valid's auc: 0.988301\n[70]    train's auc: 0.982942   valid's auc: 0.988885\n[80]    train's auc: 0.983197   valid's auc: 0.989289\n[90]    train's auc: 0.983447   valid's auc: 0.989419\n[100]   train's auc: 0.983683   valid's auc: 0.989379\n[110]   train's auc: 0.98386    valid's auc: 0.989594\n[120]   train's auc: 0.983984   valid's auc: 0.989651\n[130]   train's auc: 0.984101   valid's auc: 0.989784\n[140]   train's auc: 0.984219   valid's auc: 0.989845\n[150]   train's auc: 0.984317   valid's auc: 0.989851\n[160]   train's auc: 0.984398   valid's auc: 0.990049\n[170]   train's auc: 0.984523   valid's auc: 0.99027\n[180]   train's auc: 0.984617   valid's auc: 0.990348\n[190]   train's auc: 0.984713   valid's auc: 0.990424\n[200]   train's auc: 0.984799   valid's auc: 0.990447\n[210]   train's auc: 0.984853   valid's auc: 0.990475\n[220]   train's auc: 0.984915   valid's auc: 0.990588\n[230]   train's auc: 0.984979   valid's auc: 0.990649\n[240]   train's auc: 0.985035   valid's auc: 0.990668\n[250]   train's auc: 0.985089   valid's auc: 0.990664\n[260]   train's auc: 0.985126   valid's auc: 0.990703\n[270]   train's auc: 0.985166   valid's auc: 0.990704\n[280]   train's auc: 0.985209   valid's auc: 0.990698\n[290]   train's auc: 0.985251   valid's auc: 0.990775\n[300]   train's auc: 0.985285   valid's auc: 0.990797\n[310]   train's auc: 0.985315   valid's auc: 0.99079\n[320]   train's auc: 0.985351   valid's auc: 0.990758\n[330]   train's auc: 0.985379   valid's auc: 0.990737\n[340]   train's auc: 0.985418   valid's auc: 0.990764\n[350]   train's auc: 0.98545    valid's auc: 0.99076\nEarly stopping, best iteration is:\n[307]   train's auc: 0.985306   valid's auc: 0.990801\n\nUsing 32GB Machine with 100GB Swap Space on 2 SSD's. I usually starts the script at night and wake up in the morning to see the beautiful results :)\n\nPS: 7 Features gave me a LB score of 0.9788.\n\nThe next step is to add 3 more new features and see if it can get a score of 0.98XX",
    "317674": "@Samrat\nThanks for the share.\n50 rounds?\nI use 20, do you find cases where after 20 rounds of no better AUC suddenly it jumps up again?\n\n+ It's amusing that your Validation AUC is higher than the Training AUC, Don't you prefer to take a specific planned set of data instead of just the tail? Like day9 \\ day8 \\ specific hours?",
    "319979": "I think that you have very small valid set, and delta between valid score and LB score is really very big...",
    "323006": "Ensemble with another model of mine gives LB score of 0.9806.\n\nThat's impressive. ",
    "322974": "**Update - May 04th** \n\n- Added 6 more features and removed few under performing features (28 new features in total) and the local validation score 0.99157 and the LB score improved to 0.9803.\nEnsemble with another model of mine gives LB score of 0.9806.\n\nGiving one last try..... adding 2 new features and removed 3 features",
    "322172": "hi, could you tell me how did you generate your local cv dataset?",
    "319886": "hello,thanks for your sharing.How much memory did you use? Did you read csv in chunk or read them all ? When I read them all together, I got a memoryerr exception,but    the total memory is 120GB.I am really comfused.",
    "319715": "**Update** - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.",
    "319232": "**Update** - Added 7 more features (20 new features in total) and the local validation score jumped to 0.991275 but sadly the LB score improved just to 0.9801.",
    "318537": "9 new features, 0.9788",
    "318464": "Thanks for the share, I am wondering if &lt; 0.98XX is the limit of single model and you just give the answer.",
    "318010": "and this my first competitions too :)",
    "318009": "I use 100 million data, got 0.9791, 9 Features ; I think use more data get higher score。 but I dont have more memory in my laptop",
    "317615": "PS: 7 Features gave me a LB score of 0.9788.\n\nIs the model using only 7 features to get 0.9788 ?\nThanks .",
    "317294": "Thanks for the share."
  }
}