{
  "id": 56272,
  "title": "Small trick to iterate much faster... ",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56272",
  "author_name": "",
  "post_date": "2018-05-08T07:13:51.824298900Z",
  "votes": 27,
  "comment_count": 15,
  "views": 0,
  "content": "<p>There was a trick to avoid using huge amounts of data: downsampling only negative examples. I could reach 0.9810 on public LB with a single model trained on 900K rows (450K positive examples, 450K negative). When I compared the results with those from training on the whole data, they were the same or even better, so I could iterate much faster. At least this worked on my score level.</p>\n\n<p>EDIT: Of course I needed to calculate features on the whole dataset. But I used Google BigQuery for this, it is awesome! </p>",
  "messages": [
    {
      "id": "325151",
      "postDate": "05/08/2018 07:13:51",
      "content": "<p>There was a trick to avoid using huge amounts of data: downsampling only negative examples. I could reach 0.9810 on public LB with a single model trained on 900K rows (450K positive examples, 450K negative). When I compared the results with those from training on the whole data, they were the same or even better, so I could iterate much faster. At least this worked on my score level.</p>\n\n<p>EDIT: Of course I needed to calculate features on the whole dataset. But I used Google BigQuery for this, it is awesome! </p>",
      "rawMarkdown": "There was a trick to avoid using huge amounts of data: downsampling only negative examples. I could reach 0.9810 on public LB with a single model trained on 900K rows (450K positive examples, 450K negative). When I compared the results with those from training on the whole data, they were the same or even better, so I could iterate much faster. At least this worked on my score level.\n\nEDIT: Of course I needed to calculate features on the whole dataset. But I used Google BigQuery for this, it is awesome!",
      "votes": null
    },
    {
      "id": "325161",
      "postDate": "05/08/2018 07:21:15",
      "content": "<p>Nice trick, thx for sharing! </p>",
      "rawMarkdown": "Nice trick, thx for sharing!",
      "votes": null
    },
    {
      "id": "325306",
      "postDate": "05/08/2018 09:40:09",
      "content": "<p>You're welcome :)</p>",
      "rawMarkdown": "You're welcome :)",
      "votes": null
    },
    {
      "id": "325499",
      "postDate": "05/08/2018 13:27:38",
      "content": "<p>That’s interesting, we considered up sampling but missed down sampling :) only 900K rows? That’s amazing. Lessons learnt! Thank you very much! </p>",
      "rawMarkdown": "That’s interesting, we considered up sampling but missed down sampling :) only 900K rows? That’s amazing. Lessons learnt! Thank you very much!",
      "votes": null
    },
    {
      "id": "325579",
      "postDate": "05/08/2018 15:28:06",
      "content": "<p>Are you sampling equally from days, hours?</p>",
      "rawMarkdown": "Are you sampling equally from days, hours?",
      "votes": null
    },
    {
      "id": "325717",
      "postDate": "05/08/2018 19:31:31",
      "content": "<p>Did you do feature engineering before the downsampling?</p>",
      "rawMarkdown": "Did you do feature engineering before the downsampling?",
      "votes": null
    },
    {
      "id": "325721",
      "postDate": "05/08/2018 19:35:36",
      "content": "<p>No, I was just taking a random sample from all negative examples.</p>",
      "rawMarkdown": "No, I was just taking a random sample from all negative examples.",
      "votes": null
    },
    {
      "id": "325722",
      "postDate": "05/08/2018 19:36:24",
      "content": "<p>Yes, as I said already :)</p>",
      "rawMarkdown": "Yes, as I said already :)",
      "votes": null
    },
    {
      "id": "325728",
      "postDate": "05/08/2018 19:44:29",
      "content": "<p>I'm sorry! I read your thread earlier and I haven't seen your edit. Thanks for sharing and sorry again!</p>",
      "rawMarkdown": "I'm sorry! I read your thread earlier and I haven't seen your edit. Thanks for sharing and sorry again!",
      "votes": null
    },
    {
      "id": "325730",
      "postDate": "05/08/2018 19:45:17",
      "content": "<p>No problem :)</p>",
      "rawMarkdown": "No problem :)",
      "votes": null
    },
    {
      "id": "325776",
      "postDate": "05/08/2018 21:04:06",
      "content": "<p>It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful. Thanks Kamil for sharing this!</p>",
      "rawMarkdown": "It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful. Thanks Kamil for sharing this!",
      "votes": null
    },
    {
      "id": "325907",
      "postDate": "05/09/2018 03:23:20",
      "content": "<p>Amazing and very interesting. Thank you!</p>",
      "rawMarkdown": "Amazing and very interesting. Thank you!",
      "votes": null
    },
    {
      "id": "327314",
      "postDate": "05/11/2018 08:54:56",
      "content": "<blockquote>\n  <p>\"downsampling only negative examples\"</p>\n</blockquote>\n\n<p>Yes this is called <a href=\"https://en.wikipedia.org/wiki/Stratified_sampling\"><strong>Stratified Sampling</strong></a>. It's well-known.</p>\n\n<blockquote>\n  <p><strong>eric wrote</strong></p>\n  \n  <blockquote>\n    <p>It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful.</p>\n  </blockquote>\n</blockquote>\n\n<p>It depends entirely on the type of model. Some models are more insensitive to class imbalance than others. Some allow individual weights, e.g. you could weight the majority class examples with 1/N to reduce their influence.</p>\n\n<p>But apart from that, operationally speaking, downsampling to ~1:1 class ratio (after feature extraction, of course) reduces the amount of data and makes training much faster.</p>",
      "rawMarkdown": "&gt; \"downsampling only negative examples\"\n\nYes this is called [**Stratified Sampling**](https://en.wikipedia.org/wiki/Stratified_sampling). It's well-known.\n\n&gt; **eric wrote**\n&gt; \n&gt; &gt; It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful.\n\nIt depends entirely on the type of model. Some models are more insensitive to class imbalance than others. Some allow individual weights, e.g. you could weight the majority class examples with 1/N to reduce their influence.\n\nBut apart from that, operationally speaking, downsampling to ~1:1 class ratio (after feature extraction, of course) reduces the amount of data and makes training much faster.",
      "votes": null
    },
    {
      "id": "327318",
      "postDate": "05/11/2018 09:10:05",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "327325",
      "postDate": "05/11/2018 09:30:27",
      "content": "<p>Nice to hear that stratified sampling plays such a role. Thks for the feedback. I am upvoting for you :-)</p>",
      "rawMarkdown": "Nice to hear that stratified sampling plays such a role. Thks for the feedback. I am upvoting for you :-)",
      "votes": null
    },
    {
      "id": "327876",
      "postDate": "05/12/2018 19:56:32",
      "content": "<p>Thanks Kamil!</p>",
      "rawMarkdown": "Thanks Kamil!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325161,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "05/08/2018 07:21:15",
      "content": "<p>Nice trick, thx for sharing! </p>",
      "votes": null,
      "replies": [
        {
          "id": 325306,
          "author_name": "kamilkk",
          "author_url": "",
          "post_date": "05/08/2018 09:40:09",
          "content": "<p>You're welcome :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325499,
      "author_name": "strider1125",
      "author_url": "",
      "post_date": "05/08/2018 13:27:38",
      "content": "<p>That’s interesting, we considered up sampling but missed down sampling :) only 900K rows? That’s amazing. Lessons learnt! Thank you very much! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325579,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "05/08/2018 15:28:06",
      "content": "<p>Are you sampling equally from days, hours?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325721,
          "author_name": "kamilkk",
          "author_url": "",
          "post_date": "05/08/2018 19:35:36",
          "content": "<p>No, I was just taking a random sample from all negative examples.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325717,
      "author_name": "paulofelipe",
      "author_url": "",
      "post_date": "05/08/2018 19:31:31",
      "content": "<p>Did you do feature engineering before the downsampling?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325722,
          "author_name": "kamilkk",
          "author_url": "",
          "post_date": "05/08/2018 19:36:24",
          "content": "<p>Yes, as I said already :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325728,
          "author_name": "paulofelipe",
          "author_url": "",
          "post_date": "05/08/2018 19:44:29",
          "content": "<p>I'm sorry! I read your thread earlier and I haven't seen your edit. Thanks for sharing and sorry again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325730,
          "author_name": "kamilkk",
          "author_url": "",
          "post_date": "05/08/2018 19:45:17",
          "content": "<p>No problem :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325776,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/08/2018 21:04:06",
      "content": "<p>It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful. Thanks Kamil for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325907,
      "author_name": "closest",
      "author_url": "",
      "post_date": "05/09/2018 03:23:20",
      "content": "<p>Amazing and very interesting. Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327314,
      "author_name": "smcinerney",
      "author_url": "",
      "post_date": "05/11/2018 08:54:56",
      "content": "<blockquote>\n  <p>\"downsampling only negative examples\"</p>\n</blockquote>\n\n<p>Yes this is called <a href=\"https://en.wikipedia.org/wiki/Stratified_sampling\"><strong>Stratified Sampling</strong></a>. It's well-known.</p>\n\n<blockquote>\n  <p><strong>eric wrote</strong></p>\n  \n  <blockquote>\n    <p>It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful.</p>\n  </blockquote>\n</blockquote>\n\n<p>It depends entirely on the type of model. Some models are more insensitive to class imbalance than others. Some allow individual weights, e.g. you could weight the majority class examples with 1/N to reduce their influence.</p>\n\n<p>But apart from that, operationally speaking, downsampling to ~1:1 class ratio (after feature extraction, of course) reduces the amount of data and makes training much faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 327325,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/11/2018 09:30:27",
          "content": "<p>Nice to hear that stratified sampling plays such a role. Thks for the feedback. I am upvoting for you :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327318,
      "author_name": "danielhsu719",
      "author_url": "",
      "post_date": "05/11/2018 09:10:05",
      "content": "<p>Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327876,
      "author_name": "rohitkmgupta",
      "author_url": "",
      "post_date": "05/12/2018 19:56:32",
      "content": "<p>Thanks Kamil!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325151": "There was a trick to avoid using huge amounts of data: downsampling only negative examples. I could reach 0.9810 on public LB with a single model trained on 900K rows (450K positive examples, 450K negative). When I compared the results with those from training on the whole data, they were the same or even better, so I could iterate much faster. At least this worked on my score level.\n\nEDIT: Of course I needed to calculate features on the whole dataset. But I used Google BigQuery for this, it is awesome!",
    "325161": "Nice trick, thx for sharing!",
    "325306": "You're welcome :)",
    "325499": "That’s interesting, we considered up sampling but missed down sampling :) only 900K rows? That’s amazing. Lessons learnt! Thank you very much!",
    "325579": "Are you sampling equally from days, hours?",
    "325717": "Did you do feature engineering before the downsampling?",
    "325721": "No, I was just taking a random sample from all negative examples.",
    "325722": "Yes, as I said already :)",
    "325728": "I'm sorry! I read your thread earlier and I haven't seen your edit. Thanks for sharing and sorry again!",
    "325730": "No problem :)",
    "325776": "It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful. Thanks Kamil for sharing this!",
    "325907": "Amazing and very interesting. Thank you!",
    "327314": "&gt; \"downsampling only negative examples\"\n\nYes this is called [**Stratified Sampling**](https://en.wikipedia.org/wiki/Stratified_sampling). It's well-known.\n\n&gt; **eric wrote**\n&gt; \n&gt; &gt; It is interesting to learn that putting in a sense equal weight to the two categories makes the training easier for the model. If sample is out of balance, this trick can be useful.\n\nIt depends entirely on the type of model. Some models are more insensitive to class imbalance than others. Some allow individual weights, e.g. you could weight the majority class examples with 1/N to reduce their influence.\n\nBut apart from that, operationally speaking, downsampling to ~1:1 class ratio (after feature extraction, of course) reduces the amount of data and makes training much faster.",
    "327318": "Thank you!",
    "327325": "Nice to hear that stratified sampling plays such a role. Thks for the feedback. I am upvoting for you :-)",
    "327876": "Thanks Kamil!"
  },
  "source": "meta"
}