{
  "id": 52927,
  "title": "From 100k to 100MM rows",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52927",
  "author_name": "",
  "post_date": "2018-03-24T22:28:18.671089700Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>2 interesting facts about the following basic model</p>\n\n<p><strong>Model</strong> _ H2O random forest (ntrees=200, stopping_rounds=2, score_each_iteration=True) <br>\n<strong>Predictors</strong> _ ip, app, device, os, channel</p>\n\n<ol>\n<li><p>The importance of the feature <strong>app</strong> increases with the increase in the size of the training set (shown in the attached screenshot of the figures).</p></li>\n<li><p>Leaderboard score increased rapidly for <em>100k</em> training rows to <em>10MM</em> training rows.  The jump from <em>10MM</em> rows to <em>100MM</em> was less significant. The training time of 10MM-row training set was 9 minutes on 20 cores, while the training time of 100MM-row training set was 43 minutes.  That's why I may stick to the 10MM​ rows for now.</p>\n\n<pre><code>   0.8298 _ random sample of 100,000 rows,\n   0.9023 _ random sample of 1,000,000 rows,\n   0.9416 _ random sample of 10,000,000 rows, \n   0.9496 _ random sample of 100,000,000 rows.\n</code></pre></li>\n</ol>",
  "messages": [
    {
      "id": "302877",
      "postDate": "03/24/2018 22:28:18",
      "content": "<p>2 interesting facts about the following basic model</p>\n\n<p><strong>Model</strong> _ H2O random forest (ntrees=200, stopping_rounds=2, score_each_iteration=True) <br>\n<strong>Predictors</strong> _ ip, app, device, os, channel</p>\n\n<ol>\n<li><p>The importance of the feature <strong>app</strong> increases with the increase in the size of the training set (shown in the attached screenshot of the figures).</p></li>\n<li><p>Leaderboard score increased rapidly for <em>100k</em> training rows to <em>10MM</em> training rows.  The jump from <em>10MM</em> rows to <em>100MM</em> was less significant. The training time of 10MM-row training set was 9 minutes on 20 cores, while the training time of 100MM-row training set was 43 minutes.  That's why I may stick to the 10MM​ rows for now.</p>\n\n<pre><code>   0.8298 _ random sample of 100,000 rows,\n   0.9023 _ random sample of 1,000,000 rows,\n   0.9416 _ random sample of 10,000,000 rows, \n   0.9496 _ random sample of 100,000,000 rows.\n</code></pre></li>\n</ol>",
      "rawMarkdown": "2 interesting facts about the following basic model\n\n**Model** _ H2O random forest (ntrees=200, stopping_rounds=2, score_each_iteration=True)    \n**Predictors** _ ip, app, device, os, channel\n\n 1. The importance of the feature **app** increases with the increase in the size of the training set (shown in the attached screenshot of the figures).\n\n 2. Leaderboard score increased rapidly for *100k* training rows to *10MM* training rows.  The jump from *10MM* rows to *100MM* was less significant. The training time of 10MM-row training set was 9 minutes on 20 cores, while the training time of 100MM-row training set was 43 minutes.  That's why I may stick to the 10MM​ rows for now.\n\n           0.8298 _ random sample of 100,000 rows,\n           0.9023 _ random sample of 1,000,000 rows,\n           0.9416 _ random sample of 10,000,000 rows, \n           0.9496 _ random sample of 100,000,000 rows.",
      "votes": null
    },
    {
      "id": "302902",
      "postDate": "03/25/2018 00:10:05",
      "content": "<p>Thanks for sharing your insights. I'll take some leads from here, However, don't you think inference based on LB score can lead to not so informed decision, would you mind sharing your inferences based on your local CV?</p>",
      "rawMarkdown": "Thanks for sharing your insights. I'll take some leads from here, However, don't you think inference based on LB score can lead to not so informed decision, would you mind sharing your inferences based on your local CV?",
      "votes": null
    },
    {
      "id": "302914",
      "postDate": "03/25/2018 01:11:21",
      "content": "<p>Sure, here are my scores based on the local test set</p>\n\n<pre><code>  0.8137 _ random sample of 100,000 rows,\n  0.9251 _ random sample of 1,000,000 rows,\n  0.9558 _ random sample of 10,000,000 rows,\n  0.9647 _ random sample of 100,000,000 rows.\n</code></pre>\n\n<p>Well, honestly, I just realized that I didn't specify that the actual training happened on 40% of the data. <br>\nThus, </p>\n\n<pre><code>  0.8137 _ trained on a random sample of 60,000 rows,\n  0.9251 _ trained on a random sample of 600,000 rows,\n  0.9558 _ trained on a random sample of 6,000,000 rows,\n  0.9647 _ trained on a random sample of 60,000,000 rows.\n</code></pre>\n\n<p>This is just an unpolished experiment with random train-test splits. I am yet to think about validation. </p>",
      "rawMarkdown": "Sure, here are my scores based on the local test set\n\n      0.8137 _ random sample of 100,000 rows,\n      0.9251 _ random sample of 1,000,000 rows,\n      0.9558 _ random sample of 10,000,000 rows,\n      0.9647 _ random sample of 100,000,000 rows.\n\nWell, honestly, I just realized that I didn't specify that the actual training happened on 40% of the data.    \nThus, \n\n      0.8137 _ trained on a random sample of 60,000 rows,\n      0.9251 _ trained on a random sample of 600,000 rows,\n      0.9558 _ trained on a random sample of 6,000,000 rows,\n      0.9647 _ trained on a random sample of 60,000,000 rows.\n\nThis is just an unpolished experiment with random train-test splits. I am yet to think about validation.",
      "votes": null
    },
    {
      "id": "302920",
      "postDate": "03/25/2018 01:40:46",
      "content": "<p>Thanks for sharing, I do see what you concluded. For now, I am mostly interested in data towards the end of the time series. \nOn a different note I would assume your test train split is set to <code>shuffle=False</code>  ?</p>",
      "rawMarkdown": "Thanks for sharing, I do see what you concluded. For now, I am mostly interested in data towards the end of the time series. \nOn a different note I would assume your test train split is set to `shuffle=False`  ?",
      "votes": null
    },
    {
      "id": "303304",
      "postDate": "03/26/2018 01:14:02",
      "content": "<p>Perfectly noted, apparently, I didn't shuffle the data. </p>\n\n<p>I did the splitting with <em>h2o.splitFrame</em> and now I wonder why there is an argument \"seed\". But I did an initial search and it seems that <em>h2o.splitFrame</em> doesn't shuffle the data.</p>",
      "rawMarkdown": "Perfectly noted, apparently, I didn't shuffle the data. \n\nI did the splitting with *h2o.splitFrame* and now I wonder why there is an argument \"seed\". But I did an initial search and it seems that *h2o.splitFrame* doesn't shuffle the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 302902,
      "author_name": "shivrajp",
      "author_url": "",
      "post_date": "03/25/2018 00:10:05",
      "content": "<p>Thanks for sharing your insights. I'll take some leads from here, However, don't you think inference based on LB score can lead to not so informed decision, would you mind sharing your inferences based on your local CV?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302914,
      "author_name": "araksstepanyan",
      "author_url": "",
      "post_date": "03/25/2018 01:11:21",
      "content": "<p>Sure, here are my scores based on the local test set</p>\n\n<pre><code>  0.8137 _ random sample of 100,000 rows,\n  0.9251 _ random sample of 1,000,000 rows,\n  0.9558 _ random sample of 10,000,000 rows,\n  0.9647 _ random sample of 100,000,000 rows.\n</code></pre>\n\n<p>Well, honestly, I just realized that I didn't specify that the actual training happened on 40% of the data. <br>\nThus, </p>\n\n<pre><code>  0.8137 _ trained on a random sample of 60,000 rows,\n  0.9251 _ trained on a random sample of 600,000 rows,\n  0.9558 _ trained on a random sample of 6,000,000 rows,\n  0.9647 _ trained on a random sample of 60,000,000 rows.\n</code></pre>\n\n<p>This is just an unpolished experiment with random train-test splits. I am yet to think about validation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 302920,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "03/25/2018 01:40:46",
          "content": "<p>Thanks for sharing, I do see what you concluded. For now, I am mostly interested in data towards the end of the time series. \nOn a different note I would assume your test train split is set to <code>shuffle=False</code>  ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303304,
          "author_name": "araksstepanyan",
          "author_url": "",
          "post_date": "03/26/2018 01:14:02",
          "content": "<p>Perfectly noted, apparently, I didn't shuffle the data. </p>\n\n<p>I did the splitting with <em>h2o.splitFrame</em> and now I wonder why there is an argument \"seed\". But I did an initial search and it seems that <em>h2o.splitFrame</em> doesn't shuffle the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "302877": "2 interesting facts about the following basic model\n\n**Model** _ H2O random forest (ntrees=200, stopping_rounds=2, score_each_iteration=True)    \n**Predictors** _ ip, app, device, os, channel\n\n 1. The importance of the feature **app** increases with the increase in the size of the training set (shown in the attached screenshot of the figures).\n\n 2. Leaderboard score increased rapidly for *100k* training rows to *10MM* training rows.  The jump from *10MM* rows to *100MM* was less significant. The training time of 10MM-row training set was 9 minutes on 20 cores, while the training time of 100MM-row training set was 43 minutes.  That's why I may stick to the 10MM​ rows for now.\n\n           0.8298 _ random sample of 100,000 rows,\n           0.9023 _ random sample of 1,000,000 rows,\n           0.9416 _ random sample of 10,000,000 rows, \n           0.9496 _ random sample of 100,000,000 rows.",
    "302902": "Thanks for sharing your insights. I'll take some leads from here, However, don't you think inference based on LB score can lead to not so informed decision, would you mind sharing your inferences based on your local CV?",
    "302914": "Sure, here are my scores based on the local test set\n\n      0.8137 _ random sample of 100,000 rows,\n      0.9251 _ random sample of 1,000,000 rows,\n      0.9558 _ random sample of 10,000,000 rows,\n      0.9647 _ random sample of 100,000,000 rows.\n\nWell, honestly, I just realized that I didn't specify that the actual training happened on 40% of the data.    \nThus, \n\n      0.8137 _ trained on a random sample of 60,000 rows,\n      0.9251 _ trained on a random sample of 600,000 rows,\n      0.9558 _ trained on a random sample of 6,000,000 rows,\n      0.9647 _ trained on a random sample of 60,000,000 rows.\n\nThis is just an unpolished experiment with random train-test splits. I am yet to think about validation.",
    "302920": "Thanks for sharing, I do see what you concluded. For now, I am mostly interested in data towards the end of the time series. \nOn a different note I would assume your test train split is set to `shuffle=False`  ?",
    "303304": "Perfectly noted, apparently, I didn't shuffle the data. \n\nI did the splitting with *h2o.splitFrame* and now I wonder why there is an argument \"seed\". But I did an initial search and it seems that *h2o.splitFrame* doesn't shuffle the data."
  },
  "source": "meta"
}