{
  "id": 347643,
  "title": "Speculation time! Why did I jump from 402 to 49th?",
  "url": "/competitions/amex-default-prediction/discussion/347643",
  "author_name": "",
  "post_date": "2022-08-25T00:32:06.695029600Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm not an expert data scientist, but really enjoyed trying a few things this competition. </p>\n<p>One of the things I tried didn't really help much on the public LB, but might(?) have been <em>a</em> way to handle the difficulties normalizing the data taken from different time periods and different seasons. </p>\n<p>In short, inspired by raddar's days overdue notebook, I created two features (which became 32 after engineering) on the entire 15 million row dataset (using 280 features to predict). One to predict the chance that the days overdue would go unpaid next statement (exact logic was a bit complex, I'll look it up later). Only 1/60 rows had this issue. Another lesser feature to predict if days overdue would be 0 next month. </p>\n<p>All my submissions using this leapfrogged my other submissions. I think it might be because I had one mega feature that got TRAINED on the private dataset. With enough logical correlation to the unknown true target that it helped my model partially avoid the challenges of predicting from a different population than was trained on?</p>\n<p>What do you think?</p>",
  "messages": [
    {
      "id": "1912745",
      "postDate": "08/25/2022 00:32:06",
      "content": "<p>I'm not an expert data scientist, but really enjoyed trying a few things this competition. </p>\n<p>One of the things I tried didn't really help much on the public LB, but might(?) have been <em>a</em> way to handle the difficulties normalizing the data taken from different time periods and different seasons. </p>\n<p>In short, inspired by raddar's days overdue notebook, I created two features (which became 32 after engineering) on the entire 15 million row dataset (using 280 features to predict). One to predict the chance that the days overdue would go unpaid next statement (exact logic was a bit complex, I'll look it up later). Only 1/60 rows had this issue. Another lesser feature to predict if days overdue would be 0 next month. </p>\n<p>All my submissions using this leapfrogged my other submissions. I think it might be because I had one mega feature that got TRAINED on the private dataset. With enough logical correlation to the unknown true target that it helped my model partially avoid the challenges of predicting from a different population than was trained on?</p>\n<p>What do you think?</p>",
      "rawMarkdown": "I'm not an expert data scientist, but really enjoyed trying a few things this competition. \n\nOne of the things I tried didn't really help much on the public LB, but might(?) have been *a* way to handle the difficulties normalizing the data taken from different time periods and different seasons. \n\nIn short, inspired by raddar's days overdue notebook, I created two features (which became 32 after engineering) on the entire 15 million row dataset (using 280 features to predict). One to predict the chance that the days overdue would go unpaid next statement (exact logic was a bit complex, I'll look it up later). Only 1/60 rows had this issue. Another lesser feature to predict if days overdue would be 0 next month. \n\nAll my submissions using this leapfrogged my other submissions. I think it might be because I had one mega feature that got TRAINED on the private dataset. With enough logical correlation to the unknown true target that it helped my model partially avoid the challenges of predicting from a different population than was trained on?\n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "1912748",
      "postDate": "08/25/2022 00:34:13",
      "content": "<p>I also left off B_29 from the very start. Maybe that was the better choice this time? Not in any model (probably should've had it in the days overdue model but missed it then didn't have time to fix later)</p>",
      "rawMarkdown": "I also left off B_29 from the very start. Maybe that was the better choice this time? Not in any model (probably should've had it in the days overdue model but missed it then didn't have time to fix later)",
      "votes": null
    },
    {
      "id": "1912791",
      "postDate": "08/25/2022 01:07:08",
      "content": "<p>The days overdue label code I used was also tricky to decide on, as it wasn't clear from manual inspection the rhyme or reason of days between statements, days overdue increase vs total days increase, it wasn't as black and white as I expected.</p>\n<p>The exact labeling logic is really all in this line of code: If D39 went from 60 to 75, it would be \"-15\" in the D_39_lag value. And S_2 was in days, so if S_2_lag was -15, then there were only 15 days between statements.</p>\n<pre><code>test.loc[ (test[\"D_39_lag\"] &lt;= -28) | ((test[\"D_39_lag\"] &lt; -14) &amp; (test[\"S_2_lag\"] &gt;= test[\"D_39_lag\"]) &amp; (test[\"D_39\"] - test[\"D_39_lag\"] &gt;= 28)), [\"miss_next_payment\"] ] = 1\n</code></pre>",
      "rawMarkdown": "The days overdue label code I used was also tricky to decide on, as it wasn't clear from manual inspection the rhyme or reason of days between statements, days overdue increase vs total days increase, it wasn't as black and white as I expected.\n\nThe exact labeling logic is really all in this line of code: If D39 went from 60 to 75, it would be \"-15\" in the D_39_lag value. And S_2 was in days, so if S_2_lag was -15, then there were only 15 days between statements.\n\n    test.loc[ (test[\"D_39_lag\"] <= -28) | ((test[\"D_39_lag\"] < -14) & (test[\"S_2_lag\"] >= test[\"D_39_lag\"]) & (test[\"D_39\"] - test[\"D_39_lag\"] >= 28)), [\"miss_next_payment\"] ] = 1",
      "votes": null
    },
    {
      "id": "1912874",
      "postDate": "08/25/2022 02:38:39",
      "content": "<p>Congratulations Robert! This sounds like a great feature. I think this competition was about predicting the future. So any features or techniques that helped predict the future boosted models.</p>",
      "rawMarkdown": "Congratulations Robert! This sounds like a great feature. I think this competition was about predicting the future. So any features or techniques that helped predict the future boosted models.",
      "votes": null
    },
    {
      "id": "1922291",
      "postDate": "09/01/2022 11:46:48",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hello @roberthatch. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1912748,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 00:34:13",
      "content": "<p>I also left off B_29 from the very start. Maybe that was the better choice this time? Not in any model (probably should've had it in the days overdue model but missed it then didn't have time to fix later)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912791,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 01:07:08",
      "content": "<p>The days overdue label code I used was also tricky to decide on, as it wasn't clear from manual inspection the rhyme or reason of days between statements, days overdue increase vs total days increase, it wasn't as black and white as I expected.</p>\n<p>The exact labeling logic is really all in this line of code: If D39 went from 60 to 75, it would be \"-15\" in the D_39_lag value. And S_2 was in days, so if S_2_lag was -15, then there were only 15 days between statements.</p>\n<pre><code>test.loc[ (test[\"D_39_lag\"] &lt;= -28) | ((test[\"D_39_lag\"] &lt; -14) &amp; (test[\"S_2_lag\"] &gt;= test[\"D_39_lag\"]) &amp; (test[\"D_39\"] - test[\"D_39_lag\"] &gt;= 28)), [\"miss_next_payment\"] ] = 1\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912874,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/25/2022 02:38:39",
      "content": "<p>Congratulations Robert! This sounds like a great feature. I think this competition was about predicting the future. So any features or techniques that helped predict the future boosted models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1922291,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 11:46:48",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912745": "I'm not an expert data scientist, but really enjoyed trying a few things this competition. \n\nOne of the things I tried didn't really help much on the public LB, but might(?) have been *a* way to handle the difficulties normalizing the data taken from different time periods and different seasons. \n\nIn short, inspired by raddar's days overdue notebook, I created two features (which became 32 after engineering) on the entire 15 million row dataset (using 280 features to predict). One to predict the chance that the days overdue would go unpaid next statement (exact logic was a bit complex, I'll look it up later). Only 1/60 rows had this issue. Another lesser feature to predict if days overdue would be 0 next month. \n\nAll my submissions using this leapfrogged my other submissions. I think it might be because I had one mega feature that got TRAINED on the private dataset. With enough logical correlation to the unknown true target that it helped my model partially avoid the challenges of predicting from a different population than was trained on?\n\nWhat do you think?",
    "1912748": "I also left off B_29 from the very start. Maybe that was the better choice this time? Not in any model (probably should've had it in the days overdue model but missed it then didn't have time to fix later)",
    "1912791": "The days overdue label code I used was also tricky to decide on, as it wasn't clear from manual inspection the rhyme or reason of days between statements, days overdue increase vs total days increase, it wasn't as black and white as I expected.\n\nThe exact labeling logic is really all in this line of code: If D39 went from 60 to 75, it would be \"-15\" in the D_39_lag value. And S_2 was in days, so if S_2_lag was -15, then there were only 15 days between statements.\n\n    test.loc[ (test[\"D_39_lag\"] <= -28) | ((test[\"D_39_lag\"] < -14) & (test[\"S_2_lag\"] >= test[\"D_39_lag\"]) & (test[\"D_39\"] - test[\"D_39_lag\"] >= 28)), [\"miss_next_payment\"] ] = 1",
    "1912874": "Congratulations Robert! This sounds like a great feature. I think this competition was about predicting the future. So any features or techniques that helped predict the future boosted models.",
    "1922291": "Hello @roberthatch. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
  },
  "source": "meta"
}