{
  "id": 56339,
  "title": "The Biggest Failure",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56339",
  "author_name": "",
  "post_date": "2018-05-08T19:15:34.080984600Z",
  "votes": 28,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Like all of you, I had a very emotional night yesterday albeit for a slightly different reason:</p>\n\n<p><img src=\"https://s31.postimg.cc/ykv56sl8r/Screen_Shot_2018-05-08_at_1.24.07_PM.png\" alt=\"enter image description here\"></p>\n\n<p>Having perused through the entire LB, I actually experienced the single largest drop out of any contestant in this competition (and any of my previous competitions) and was the only person to experience a 4-digit LB drop period. Having invested months into the competition + some new hardware, this was grueling. In all, it was a very humbling and particularly embarrassing public experience that will undoubtedly mark my Kaggle history.</p>\n\n<p>Having looked through my private LB submissions, in hindsight it's clear that things started going sour as soon as I introduced target encoded features after crossing the 0.9812083 private /\n0.9807667 public line. Looking through my code, I had absentmindedly introduced leakage into my data pre-processing pipeline. I understand how the leakage bolstered my local validation; but what still surprises me is how said leakage caused severe overfitting against the public test set (public LB). When I saw ~9813 local val across multiple models and 9814 on public LB, I felt confident I was on the right track and did not bother to triple check my processing code. I'm saying this publicly specifically because I had another thread on speeding up lightgbm convergence via LR decay, and didn't want anyone to refuse to experiment with that based simply upon my poor performance by associating the two. They're mutually exclusive :-).</p>\n\n<p>In my mid-competition comment on <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55788#321669\">lessons learned</a>, I had mentioned the need for developing a good representative validation set. This is doubly true for me now. Reading over some of the solutions posted by others, one technique someone had success with was downsampling negative cases to equal the positive cases. Other people selected specific hours to validate against. I think the main take away is to invest time into it and try different strategies until something works, but continue to occasionally test your predictions against other strategies so that you never fall into traps like the aforementioned one.</p>\n\n<p>Anyhow, congrats to all the winners and hard workers, and I hope to see many of you guys on Avito if you have time.</p>",
  "messages": [
    {
      "id": "325712",
      "postDate": "05/08/2018 19:15:34",
      "content": "<p>Like all of you, I had a very emotional night yesterday albeit for a slightly different reason:</p>\n\n<p><img src=\"https://s31.postimg.cc/ykv56sl8r/Screen_Shot_2018-05-08_at_1.24.07_PM.png\" alt=\"enter image description here\"></p>\n\n<p>Having perused through the entire LB, I actually experienced the single largest drop out of any contestant in this competition (and any of my previous competitions) and was the only person to experience a 4-digit LB drop period. Having invested months into the competition + some new hardware, this was grueling. In all, it was a very humbling and particularly embarrassing public experience that will undoubtedly mark my Kaggle history.</p>\n\n<p>Having looked through my private LB submissions, in hindsight it's clear that things started going sour as soon as I introduced target encoded features after crossing the 0.9812083 private /\n0.9807667 public line. Looking through my code, I had absentmindedly introduced leakage into my data pre-processing pipeline. I understand how the leakage bolstered my local validation; but what still surprises me is how said leakage caused severe overfitting against the public test set (public LB). When I saw ~9813 local val across multiple models and 9814 on public LB, I felt confident I was on the right track and did not bother to triple check my processing code. I'm saying this publicly specifically because I had another thread on speeding up lightgbm convergence via LR decay, and didn't want anyone to refuse to experiment with that based simply upon my poor performance by associating the two. They're mutually exclusive :-).</p>\n\n<p>In my mid-competition comment on <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55788#321669\">lessons learned</a>, I had mentioned the need for developing a good representative validation set. This is doubly true for me now. Reading over some of the solutions posted by others, one technique someone had success with was downsampling negative cases to equal the positive cases. Other people selected specific hours to validate against. I think the main take away is to invest time into it and try different strategies until something works, but continue to occasionally test your predictions against other strategies so that you never fall into traps like the aforementioned one.</p>\n\n<p>Anyhow, congrats to all the winners and hard workers, and I hope to see many of you guys on Avito if you have time.</p>",
      "rawMarkdown": "Like all of you, I had a very emotional night yesterday albeit for a slightly different reason:\n\n![enter image description here][1]\n\nHaving perused through the entire LB, I actually experienced the single largest drop out of any contestant in this competition (and any of my previous competitions) and was the only person to experience a 4-digit LB drop period. Having invested months into the competition + some new hardware, this was grueling. In all, it was a very humbling and particularly embarrassing public experience that will undoubtedly mark my Kaggle history.\n\nHaving looked through my private LB submissions, in hindsight it's clear that things started going sour as soon as I introduced target encoded features after crossing the 0.9812083 private /\n0.9807667 public line. Looking through my code, I had absentmindedly introduced leakage into my data pre-processing pipeline. I understand how the leakage bolstered my local validation; but what still surprises me is how said leakage caused severe overfitting against the public test set (public LB). When I saw ~9813 local val across multiple models and 9814 on public LB, I felt confident I was on the right track and did not bother to triple check my processing code. I'm saying this publicly specifically because I had another thread on speeding up lightgbm convergence via LR decay, and didn't want anyone to refuse to experiment with that based simply upon my poor performance by associating the two. They're mutually exclusive :-).\n\nIn my mid-competition comment on [lessons learned][2], I had mentioned the need for developing a good representative validation set. This is doubly true for me now. Reading over some of the solutions posted by others, one technique someone had success with was downsampling negative cases to equal the positive cases. Other people selected specific hours to validate against. I think the main take away is to invest time into it and try different strategies until something works, but continue to occasionally test your predictions against other strategies so that you never fall into traps like the aforementioned one.\n\nAnyhow, congrats to all the winners and hard workers, and I hope to see many of you guys on Avito if you have time.\n\n\n  [1]: https://s31.postimg.cc/ykv56sl8r/Screen_Shot_2018-05-08_at_1.24.07_PM.png\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55788#321669",
      "votes": null
    },
    {
      "id": "325719",
      "postDate": "05/08/2018 19:34:00",
      "content": "<p>Thank you for your share of experience, and your advice in the middle of the competition, see you on Avito :)</p>",
      "rawMarkdown": "Thank you for your share of experience, and your advice in the middle of the competition, see you on Avito :)",
      "votes": null
    },
    {
      "id": "325723",
      "postDate": "05/08/2018 19:36:28",
      "content": "<p>All we can do is learn and implement what we learned in future tasks.\nBe strong my friend, don't let it get to you!\nSee you soon in Avito =)</p>",
      "rawMarkdown": "All we can do is learn and implement what we learned in future tasks.\nBe strong my friend, don't let it get to you!\nSee you soon in Avito =)",
      "votes": null
    },
    {
      "id": "325898",
      "postDate": "05/09/2018 03:03:59",
      "content": "<p>I also experienced a significant drop in the private LB, but why history download rate hurt the model still confuse me. Do you have any idea?</p>",
      "rawMarkdown": "I also experienced a significant drop in the private LB, but why history download rate hurt the model still confuse me. Do you have any idea?",
      "votes": null
    },
    {
      "id": "325986",
      "postDate": "05/09/2018 06:38:46",
      "content": "<p>Sorry for that.  If it can make you happier I once experienced an even larger drop in LB (in my second competition I think).  That's when I learned that overfiting was a deadly poison that can be hard to detect. </p>\n\n<p>Target encoding is very dangerous, and should always be done in an out of fold manner.  And even that way it can be bad.  In this competition I could not make it work via oof, and it barely works as lag feature.  Seems other top teams experienced the same.  </p>\n\n<p>Last, thank you for pointing to the two_round parameter, without it I would not be in top 10 I think.</p>",
      "rawMarkdown": "Sorry for that.  If it can make you happier I once experienced an even larger drop in LB (in my second competition I think).  That's when I learned that overfiting was a deadly poison that can be hard to detect. \n\nTarget encoding is very dangerous, and should always be done in an out of fold manner.  And even that way it can be bad.  In this competition I could not make it work via oof, and it barely works as lag feature.  Seems other top teams experienced the same.  \n\nLast, thank you for pointing to the two_round parameter, without it I would not be in top 10 I think.",
      "votes": null
    },
    {
      "id": "326004",
      "postDate": "05/09/2018 06:50:00",
      "content": "<p>Hi authman, I am sorry to hear that. I experienced all that but in real life at school. I have been following a master in machine learning and the final project was to create a predictor using all the techniques we had seen sofar. It was at the beginning of the course and I suspect the teacher did on purpose to avoid explaining overfitting to make sure we fall in the trap... I was one of them...</p>\n\n<p>All the best to you and I hope this bad luck will not discourage you from continuing your experience at Kaggle</p>",
      "rawMarkdown": "Hi authman, I am sorry to hear that. I experienced all that but in real life at school. I have been following a master in machine learning and the final project was to create a predictor using all the techniques we had seen sofar. It was at the beginning of the course and I suspect the teacher did on purpose to avoid explaining overfitting to make sure we fall in the trap... I was one of them...\n\nAll the best to you and I hope this bad luck will not discourage you from continuing your experience at Kaggle",
      "votes": null
    },
    {
      "id": "326068",
      "postDate": "05/09/2018 08:26:56",
      "content": "<p>@CPMP\nIndeed the two_round parameter is very helpful,\nWhat is the \"SWAP\" thingy i read around?\ni could not find it in discussions now, some people say they used \"32GB RAM + 64GB Swap\" for example.. something like that.\nCould you share what it means?</p>",
      "rawMarkdown": "CPMP\nIndeed the two_round parameter is very helpful,\nWhat is the \"SWAP\" thingy i read around?\ni could not find it in discussions now, some people say they used \"32GB RAM + 64GB Swap\" for example.. something like that.\nCould you share what it means?",
      "votes": null
    },
    {
      "id": "326072",
      "postDate": "05/09/2018 08:32:36",
      "content": "<p>Swap is a way to add virtual memory to your machine.  It is built in by default in Windows, Mac OSX, and recent Ubuntu releases for instance (as well as most OS).  The best I've seen among these three is OSX.  I could run XGBoost processes up to 55 GB on a 16GB Macbook Pro without problem.  </p>\n\n<p>How to add/expand swap depends on your OS.  Check for its documentation or on stackoverflow.</p>",
      "rawMarkdown": "Swap is a way to add virtual memory to your machine.  It is built in by default in Windows, Mac OSX, and recent Ubuntu releases for instance (as well as most OS).  The best I've seen among these three is OSX.  I could run XGBoost processes up to 55 GB on a 16GB Macbook Pro without problem.  \n\nHow to add/expand swap depends on your OS.  Check for its documentation or on stackoverflow.",
      "votes": null
    },
    {
      "id": "326317",
      "postDate": "05/09/2018 14:48:46",
      "content": "<p>I feel you man. Had the same problem in toxic competition (my first competition) and blindly followed public LB, though I mitigated the leakage (convai dataset) by blending with different other models without the leak.</p>\n\n<p>For this competition, I blindly ignored public LB :) which helped me exclude target encoding from day-one and religiously tracked my local validation, though I made a terrible terrible mistake in splitting my initial dataset, too many mistakes in FE that required me to redo many times all my training again and being late in the competition. \nBut I learned a lot...</p>",
      "rawMarkdown": "I feel you man. Had the same problem in toxic competition (my first competition) and blindly followed public LB, though I mitigated the leakage (convai dataset) by blending with different other models without the leak.\n\nFor this competition, I blindly ignored public LB :) which helped me exclude target encoding from day-one and religiously tracked my local validation, though I made a terrible terrible mistake in splitting my initial dataset, too many mistakes in FE that required me to redo many times all my training again and being late in the competition. \nBut I learned a lot...",
      "votes": null
    },
    {
      "id": "326495",
      "postDate": "05/09/2018 20:15:41",
      "content": "<p>Sorry to hear that. But everyone who did not copy Dirk's kernel experienced a big failure.</p>",
      "rawMarkdown": "Sorry to hear that. But everyone who did not copy Dirk's kernel experienced a big failure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325719,
      "author_name": "marrvolo",
      "author_url": "",
      "post_date": "05/08/2018 19:34:00",
      "content": "<p>Thank you for your share of experience, and your advice in the middle of the competition, see you on Avito :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325723,
      "author_name": "tpthegreat",
      "author_url": "",
      "post_date": "05/08/2018 19:36:28",
      "content": "<p>All we can do is learn and implement what we learned in future tasks.\nBe strong my friend, don't let it get to you!\nSee you soon in Avito =)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325898,
      "author_name": "wuyhbb",
      "author_url": "",
      "post_date": "05/09/2018 03:03:59",
      "content": "<p>I also experienced a significant drop in the private LB, but why history download rate hurt the model still confuse me. Do you have any idea?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325986,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2018 06:38:46",
      "content": "<p>Sorry for that.  If it can make you happier I once experienced an even larger drop in LB (in my second competition I think).  That's when I learned that overfiting was a deadly poison that can be hard to detect. </p>\n\n<p>Target encoding is very dangerous, and should always be done in an out of fold manner.  And even that way it can be bad.  In this competition I could not make it work via oof, and it barely works as lag feature.  Seems other top teams experienced the same.  </p>\n\n<p>Last, thank you for pointing to the two_round parameter, without it I would not be in top 10 I think.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326068,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "05/09/2018 08:26:56",
          "content": "<p>@CPMP\nIndeed the two_round parameter is very helpful,\nWhat is the \"SWAP\" thingy i read around?\ni could not find it in discussions now, some people say they used \"32GB RAM + 64GB Swap\" for example.. something like that.\nCould you share what it means?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326072,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/09/2018 08:32:36",
          "content": "<p>Swap is a way to add virtual memory to your machine.  It is built in by default in Windows, Mac OSX, and recent Ubuntu releases for instance (as well as most OS).  The best I've seen among these three is OSX.  I could run XGBoost processes up to 55 GB on a 16GB Macbook Pro without problem.  </p>\n\n<p>How to add/expand swap depends on your OS.  Check for its documentation or on stackoverflow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326004,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/09/2018 06:50:00",
      "content": "<p>Hi authman, I am sorry to hear that. I experienced all that but in real life at school. I have been following a master in machine learning and the final project was to create a predictor using all the techniques we had seen sofar. It was at the beginning of the course and I suspect the teacher did on purpose to avoid explaining overfitting to make sure we fall in the trap... I was one of them...</p>\n\n<p>All the best to you and I hope this bad luck will not discourage you from continuing your experience at Kaggle</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326317,
      "author_name": "riadsouissi",
      "author_url": "",
      "post_date": "05/09/2018 14:48:46",
      "content": "<p>I feel you man. Had the same problem in toxic competition (my first competition) and blindly followed public LB, though I mitigated the leakage (convai dataset) by blending with different other models without the leak.</p>\n\n<p>For this competition, I blindly ignored public LB :) which helped me exclude target encoding from day-one and religiously tracked my local validation, though I made a terrible terrible mistake in splitting my initial dataset, too many mistakes in FE that required me to redo many times all my training again and being late in the competition. \nBut I learned a lot...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326495,
      "author_name": "kownse",
      "author_url": "",
      "post_date": "05/09/2018 20:15:41",
      "content": "<p>Sorry to hear that. But everyone who did not copy Dirk's kernel experienced a big failure.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325712": "Like all of you, I had a very emotional night yesterday albeit for a slightly different reason:\n\n![enter image description here][1]\n\nHaving perused through the entire LB, I actually experienced the single largest drop out of any contestant in this competition (and any of my previous competitions) and was the only person to experience a 4-digit LB drop period. Having invested months into the competition + some new hardware, this was grueling. In all, it was a very humbling and particularly embarrassing public experience that will undoubtedly mark my Kaggle history.\n\nHaving looked through my private LB submissions, in hindsight it's clear that things started going sour as soon as I introduced target encoded features after crossing the 0.9812083 private /\n0.9807667 public line. Looking through my code, I had absentmindedly introduced leakage into my data pre-processing pipeline. I understand how the leakage bolstered my local validation; but what still surprises me is how said leakage caused severe overfitting against the public test set (public LB). When I saw ~9813 local val across multiple models and 9814 on public LB, I felt confident I was on the right track and did not bother to triple check my processing code. I'm saying this publicly specifically because I had another thread on speeding up lightgbm convergence via LR decay, and didn't want anyone to refuse to experiment with that based simply upon my poor performance by associating the two. They're mutually exclusive :-).\n\nIn my mid-competition comment on [lessons learned][2], I had mentioned the need for developing a good representative validation set. This is doubly true for me now. Reading over some of the solutions posted by others, one technique someone had success with was downsampling negative cases to equal the positive cases. Other people selected specific hours to validate against. I think the main take away is to invest time into it and try different strategies until something works, but continue to occasionally test your predictions against other strategies so that you never fall into traps like the aforementioned one.\n\nAnyhow, congrats to all the winners and hard workers, and I hope to see many of you guys on Avito if you have time.\n\n\n  [1]: https://s31.postimg.cc/ykv56sl8r/Screen_Shot_2018-05-08_at_1.24.07_PM.png\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55788#321669",
    "325719": "Thank you for your share of experience, and your advice in the middle of the competition, see you on Avito :)",
    "325723": "All we can do is learn and implement what we learned in future tasks.\nBe strong my friend, don't let it get to you!\nSee you soon in Avito =)",
    "325898": "I also experienced a significant drop in the private LB, but why history download rate hurt the model still confuse me. Do you have any idea?",
    "325986": "Sorry for that.  If it can make you happier I once experienced an even larger drop in LB (in my second competition I think).  That's when I learned that overfiting was a deadly poison that can be hard to detect. \n\nTarget encoding is very dangerous, and should always be done in an out of fold manner.  And even that way it can be bad.  In this competition I could not make it work via oof, and it barely works as lag feature.  Seems other top teams experienced the same.  \n\nLast, thank you for pointing to the two_round parameter, without it I would not be in top 10 I think.",
    "326004": "Hi authman, I am sorry to hear that. I experienced all that but in real life at school. I have been following a master in machine learning and the final project was to create a predictor using all the techniques we had seen sofar. It was at the beginning of the course and I suspect the teacher did on purpose to avoid explaining overfitting to make sure we fall in the trap... I was one of them...\n\nAll the best to you and I hope this bad luck will not discourage you from continuing your experience at Kaggle",
    "326068": "CPMP\nIndeed the two_round parameter is very helpful,\nWhat is the \"SWAP\" thingy i read around?\ni could not find it in discussions now, some people say they used \"32GB RAM + 64GB Swap\" for example.. something like that.\nCould you share what it means?",
    "326072": "Swap is a way to add virtual memory to your machine.  It is built in by default in Windows, Mac OSX, and recent Ubuntu releases for instance (as well as most OS).  The best I've seen among these three is OSX.  I could run XGBoost processes up to 55 GB on a 16GB Macbook Pro without problem.  \n\nHow to add/expand swap depends on your OS.  Check for its documentation or on stackoverflow.",
    "326317": "I feel you man. Had the same problem in toxic competition (my first competition) and blindly followed public LB, though I mitigated the leakage (convai dataset) by blending with different other models without the leak.\n\nFor this competition, I blindly ignored public LB :) which helped me exclude target encoding from day-one and religiously tracked my local validation, though I made a terrible terrible mistake in splitting my initial dataset, too many mistakes in FE that required me to redo many times all my training again and being late in the competition. \nBut I learned a lot...",
    "326495": "Sorry to hear that. But everyone who did not copy Dirk's kernel experienced a big failure."
  },
  "source": "meta"
}