{
  "id": 93498,
  "title": "LANL Adversarial Validation - Shakeup is Coming",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93498",
  "author_name": "Bojan Tunguz",
  "post_date": "2019-05-27T15:52:30.145000",
  "votes": 17,
  "comment_count": 25,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming\">https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming</a></p>",
  "messages": [
    {
      "id": 537794,
      "postDate": "2019-05-27T15:52:30.147Z",
      "content": "<p><a href=\"https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming\">https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming</a></p>",
      "rawMarkdown": "https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming",
      "votes": 15
    },
    {
      "id": 540632,
      "postDate": "2019-05-31T19:15:19.493Z",
      "content": "<p>I've modified the code to show the feature importances. Sure enough, various \"mean\" features are the most \"dangerous\".\n<a href=\"https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming\"><img src=\"https://www.kaggleusercontent.com/kf/14989740/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..TRRcw0v2keYZBz0dkDvIlw.knDW1Jy3uRYKZdrqtyG_NTGbxKsGeJ8jhKW69_pXSWvj1Wsk3gxdimuHNyLjUwXh-JRI6n4VwWgNn6uIuEr4Daoqee13AD64BYoARHW2sMHXObhI5g27rgTBMwKD3sHmD49zwsamsUyakah84ezFHIKqKDTxGvuTwDgv6jO9Ti_8iRt1b_v-8mxqDz5Dk8xl.9sLfPTcqzxxcKn626_Bc0w/__results___files/__results___10_0.png\" alt=\"features\"></a></p>",
      "rawMarkdown": "I've modified the code to show the feature importances. Sure enough, various \"mean\" features are the most \"dangerous\".\n[![features](https://www.kaggleusercontent.com/kf/14989740/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..TRRcw0v2keYZBz0dkDvIlw.knDW1Jy3uRYKZdrqtyG_NTGbxKsGeJ8jhKW69_pXSWvj1Wsk3gxdimuHNyLjUwXh-JRI6n4VwWgNn6uIuEr4Daoqee13AD64BYoARHW2sMHXObhI5g27rgTBMwKD3sHmD49zwsamsUyakah84ezFHIKqKDTxGvuTwDgv6jO9Ti_8iRt1b_v-8mxqDz5Dk8xl.9sLfPTcqzxxcKn626_Bc0w/__results___files/__results___10_0.png)](https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming)\n",
      "votes": 3,
      "replies": [
        {
          "id": 541154,
          "postDate": "2019-06-01T22:25:41.163Z",
          "content": "<p>I think some AV difference is inevitable. My research indicates that there are more high-TTF segments in test than there are in train - another reason to be cautious about the leaderboard!</p>",
          "rawMarkdown": "I think some AV difference is inevitable. My research indicates that there are more high-TTF segments in test than there are in train - another reason to be cautious about the leaderboard!"
        },
        {
          "id": 541155,
          "postDate": "2019-06-01T22:28:29.843Z",
          "content": "<p>Your research? Please share your research. Are you going by the research paper visuals on exp. p4677?</p>",
          "rawMarkdown": "Your research? Please share your research. Are you going by the research paper visuals on exp. p4677?",
          "votes": 1
        },
        {
          "id": 541167,
          "postDate": "2019-06-01T23:59:23.343Z",
          "content": "<p>This is purely my own work using the competition data. I'll share more once the deadline is over, but I think models that cannot predict TTF &gt; 10 well (or at all) may have disappointing scores.</p>",
          "rawMarkdown": "This is purely my own work using the competition data. I'll share more once the deadline is over, but I think models that cannot predict TTF &gt; 10 well (or at all) may have disappointing scores."
        }
      ]
    },
    {
      "id": 537943,
      "postDate": "2019-05-27T22:24:50.727Z",
      "content": "<p>My adversarial validation AUC is around 0.85 as well :)</p>",
      "rawMarkdown": "My adversarial validation AUC is around 0.85 as well :)",
      "votes": 2,
      "replies": [
        {
          "id": 537949,
          "postDate": "2019-05-27T22:38:57.890Z",
          "content": "<p>So what are <strong>you</strong> doing about it?</p>",
          "rawMarkdown": "So what are **you** doing about it?",
          "votes": 2
        },
        {
          "id": 538242,
          "postDate": "2019-05-28T10:22:11.820Z",
          "content": "<p>Why do you want to do something about it?</p>",
          "rawMarkdown": "Why do you want to do something about it?",
          "votes": 2
        },
        {
          "id": 538643,
          "postDate": "2019-05-28T23:02:57.217Z",
          "content": "<p>I’m praying for the god of overfit.</p>",
          "rawMarkdown": "I’m praying for the god of overfit.",
          "votes": 7
        },
        {
          "id": 538644,
          "postDate": "2019-05-28T23:03:23.043Z",
          "content": "<p>And in the meantime selecting better features :)</p>",
          "rawMarkdown": "And in the meantime selecting better features :)",
          "votes": 4
        },
        {
          "id": 538919,
          "postDate": "2019-05-29T09:50:43.587Z",
          "content": "<p>Seriously, so you usually don't do anything about it? I put some effort to choose only the features that don't increase adversarial validation, so I have a worse CV score than I could have...</p>\n\n<p>I think it's a more general question not especially related to LANL so maybe I can count on some opinion from <a href=\"/cpmpml\">@cpmpml</a>? </p>",
          "rawMarkdown": "Seriously, so you usually don't do anything about it? I put some effort to choose only the features that don't increase adversarial validation, so I have a worse CV score than I could have...\n\nI think it's a more general question not especially related to LANL so maybe I can count on some opinion from @cpmpml? ",
          "votes": 1
        },
        {
          "id": 538921,
          "postDate": "2019-05-29T09:53:12.930Z",
          "content": "<p>All I can say is that in Malware I had an adversarial auc of 0.72 I think, yet I had great models.</p>\n\n<p>I have an adversarial auc in the same ballpark here.</p>",
          "rawMarkdown": "All I can say is that in Malware I had an adversarial auc of 0.72 I think, yet I had great models.\n\nI have an adversarial auc in the same ballpark here.",
          "votes": 2
        },
        {
          "id": 539056,
          "postDate": "2019-05-29T12:57:25.400Z",
          "content": "<p>Well, interesting. I have to think about how much 0.72 really is. And it's good to know you are observing this metric too - not only because of this competition, but generally. Thanks</p>\n\n<p>btw my adversarial auc is 0.52</p>",
          "rawMarkdown": "Well, interesting. I have to think about how much 0.72 really is. And it's good to know you are observing this metric too - not only because of this competition, but generally. Thanks\n\nbtw my adversarial auc is 0.52"
        },
        {
          "id": 539077,
          "postDate": "2019-05-29T13:36:22.477Z",
          "content": "<p>Now I understand Giba post \n&gt; And in the meantime selecting better features :)</p>\n\n<p>That he is actually trying to lower it</p>",
          "rawMarkdown": "Now I understand Giba post \n&gt; And in the meantime selecting better features :)\n\nThat he is actually trying to lower it",
          "votes": 1
        }
      ]
    },
    {
      "id": 537883,
      "postDate": "2019-05-27T19:39:43.490Z",
      "content": "<p>As others have noted, there is a significant difference in the acoustic mean between the two datasets, and I believe AV is reliant on this and other mean-weighted features. Personally, I think the best results in this competition will not rely on these features, and train/test will consequently be more similar. I could be wrong, but those are my thoughts!</p>",
      "rawMarkdown": "As others have noted, there is a significant difference in the acoustic mean between the two datasets, and I believe AV is reliant on this and other mean-weighted features. Personally, I think the best results in this competition will not rely on these features, and train/test will consequently be more similar. I could be wrong, but those are my thoughts!",
      "votes": 2
    },
    {
      "id": 537867,
      "postDate": "2019-05-27T19:10:13.897Z",
      "content": "<p>You are as addicted to AV as I am to GP ;)</p>",
      "rawMarkdown": "You are as addicted to AV as I am to GP ;)",
      "votes": 2,
      "replies": [
        {
          "id": 537878,
          "postDate": "2019-05-27T19:25:02.417Z",
          "content": "<p>You gotta have your niche. ;)</p>",
          "rawMarkdown": "You gotta have your niche. ;)",
          "votes": 4
        },
        {
          "id": 538365,
          "postDate": "2019-05-28T13:26:16.850Z",
          "content": "<p>I always appreciate the AV kernels. Keep them coming!</p>",
          "rawMarkdown": "I always appreciate the AV kernels. Keep them coming!",
          "votes": 1
        },
        {
          "id": 539252,
          "postDate": "2019-05-29T19:05:12.900Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 538965,
      "postDate": "2019-05-29T10:53:27.353Z",
      "content": "<p>I think your high AUC may just be due to curse of dimensionality. I could be wrong. How many features does it have again?</p>",
      "rawMarkdown": "I think your high AUC may just be due to curse of dimensionality. I could be wrong. How many features does it have again?",
      "replies": [
        {
          "id": 539065,
          "postDate": "2019-05-29T13:12:29.893Z",
          "content": "<p>With 16 features it's around 0.835. </p>",
          "rawMarkdown": "With 16 features it's around 0.835. ",
          "votes": 1
        },
        {
          "id": 539139,
          "postDate": "2019-05-29T15:29:05.870Z",
          "content": "<p>I'm looking at <a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a> It has well over 100 features. Is Bojan using that one or another one? I'm suggesting the high AUC for Bojan's kernel is the result of high number of features.</p>",
          "rawMarkdown": "I'm looking at https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples It has well over 100 features. Is Bojan using that one or another one? I'm suggesting the high AUC for Bojan's kernel is the result of high number of features."
        },
        {
          "id": 539182,
          "postDate": "2019-05-29T16:39:25.657Z",
          "content": "<p>I've tested different combination of features for adversarial, from 4 features to 3850 features. Best scenario was AUC around 0.64 but the downside was lower CV and LB. The balance for me is AUC around 0.84 with 25 features.  More features probably lead to higher AUC, but  this is another dilemma of this competition,  3 of my features with most importance and 0.035 effect on CV have the consequence of  AUC around 0.97. Whether to keep them or not is a question I have no answer for!</p>",
          "rawMarkdown": "I've tested different combination of features for adversarial, from 4 features to 3850 features. Best scenario was AUC around 0.64 but the downside was lower CV and LB. The balance for me is AUC around 0.84 with 25 features.  More features probably lead to higher AUC, but  this is another dilemma of this competition,  3 of my features with most importance and 0.035 effect on CV have the consequence of  AUC around 0.97. Whether to keep them or not is a question I have no answer for!",
          "votes": 1
        }
      ]
    },
    {
      "id": 537884,
      "postDate": "2019-05-27T19:40:18.960Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 537886,
          "postDate": "2019-05-27T19:41:52.403Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 537903,
          "postDate": "2019-05-27T20:07:51.617Z",
          "content": "<p>There is a significant difference between train and test sets. So it's not easy to separate them <a href=\"/zero92\">@zero92</a>. </p>",
          "rawMarkdown": "There is a significant difference between train and test sets. So it's not easy to separate them @zero92. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 540632,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2019-05-31T19:15:19.493000",
      "content": "<p>I've modified the code to show the feature importances. Sure enough, various \"mean\" features are the most \"dangerous\".\n<a href=\"https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming\"><img src=\"https://www.kaggleusercontent.com/kf/14989740/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..TRRcw0v2keYZBz0dkDvIlw.knDW1Jy3uRYKZdrqtyG_NTGbxKsGeJ8jhKW69_pXSWvj1Wsk3gxdimuHNyLjUwXh-JRI6n4VwWgNn6uIuEr4Daoqee13AD64BYoARHW2sMHXObhI5g27rgTBMwKD3sHmD49zwsamsUyakah84ezFHIKqKDTxGvuTwDgv6jO9Ti_8iRt1b_v-8mxqDz5Dk8xl.9sLfPTcqzxxcKn626_Bc0w/__results___files/__results___10_0.png\" alt=\"features\"></a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 541154,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-06-01T22:25:41.163000",
          "content": "<p>I think some AV difference is inevitable. My research indicates that there are more high-TTF segments in test than there are in train - another reason to be cautious about the leaderboard!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 541155,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-06-01T22:28:29.843000",
          "content": "<p>Your research? Please share your research. Are you going by the research paper visuals on exp. p4677?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 541167,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-06-01T23:59:23.343000",
          "content": "<p>This is purely my own work using the competition data. I'll share more once the deadline is over, but I think models that cannot predict TTF &gt; 10 well (or at all) may have disappointing scores.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 537943,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-05-27T22:24:50.727000",
      "content": "<p>My adversarial validation AUC is around 0.85 as well :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 537949,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2019-05-27T22:38:57.890000",
          "content": "<p>So what are <strong>you</strong> doing about it?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 538242,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-28T10:22:11.820000",
          "content": "<p>Why do you want to do something about it?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 538643,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-05-28T23:02:57.217000",
          "content": "<p>I’m praying for the god of overfit.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 538644,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-05-28T23:03:23.043000",
          "content": "<p>And in the meantime selecting better features :)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 538919,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-05-29T09:50:43.587000",
          "content": "<p>Seriously, so you usually don't do anything about it? I put some effort to choose only the features that don't increase adversarial validation, so I have a worse CV score than I could have...</p>\n\n<p>I think it's a more general question not especially related to LANL so maybe I can count on some opinion from <a href=\"/cpmpml\">@cpmpml</a>? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 538921,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-29T09:53:12.930000",
          "content": "<p>All I can say is that in Malware I had an adversarial auc of 0.72 I think, yet I had great models.</p>\n\n<p>I have an adversarial auc in the same ballpark here.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 539056,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-05-29T12:57:25.400000",
          "content": "<p>Well, interesting. I have to think about how much 0.72 really is. And it's good to know you are observing this metric too - not only because of this competition, but generally. Thanks</p>\n\n<p>btw my adversarial auc is 0.52</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 539077,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-05-29T13:36:22.477000",
          "content": "<p>Now I understand Giba post \n&gt; And in the meantime selecting better features :)</p>\n\n<p>That he is actually trying to lower it</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 537883,
      "author_name": "RNA",
      "author_url": "",
      "post_date": "2019-05-27T19:39:43.490000",
      "content": "<p>As others have noted, there is a significant difference in the acoustic mean between the two datasets, and I believe AV is reliant on this and other mean-weighted features. Personally, I think the best results in this competition will not rely on these features, and train/test will consequently be more similar. I could be wrong, but those are my thoughts!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 537867,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-05-27T19:10:13.897000",
      "content": "<p>You are as addicted to AV as I am to GP ;)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 537878,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2019-05-27T19:25:02.417000",
          "content": "<p>You gotta have your niche. ;)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 538365,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-05-28T13:26:16.850000",
          "content": "<p>I always appreciate the AV kernels. Keep them coming!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 539252,
          "author_name": "rafzy",
          "author_url": "",
          "post_date": "2019-05-29T19:05:12.900000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 538965,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-05-29T10:53:27.353000",
      "content": "<p>I think your high AUC may just be due to curse of dimensionality. I could be wrong. How many features does it have again?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 539065,
          "author_name": "Sia",
          "author_url": "",
          "post_date": "2019-05-29T13:12:29.893000",
          "content": "<p>With 16 features it's around 0.835. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 539139,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-05-29T15:29:05.870000",
          "content": "<p>I'm looking at <a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a> It has well over 100 features. Is Bojan using that one or another one? I'm suggesting the high AUC for Bojan's kernel is the result of high number of features.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 539182,
          "author_name": "Sia",
          "author_url": "",
          "post_date": "2019-05-29T16:39:25.657000",
          "content": "<p>I've tested different combination of features for adversarial, from 4 features to 3850 features. Best scenario was AUC around 0.64 but the downside was lower CV and LB. The balance for me is AUC around 0.84 with 25 features.  More features probably lead to higher AUC, but  this is another dilemma of this competition,  3 of my features with most importance and 0.035 effect on CV have the consequence of  AUC around 0.97. Whether to keep them or not is a question I have no answer for!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 537884,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-27T19:40:18.960000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 537886,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-27T19:41:52.403000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 537903,
          "author_name": "silverstone",
          "author_url": "",
          "post_date": "2019-05-27T20:07:51.617000",
          "content": "<p>There is a significant difference between train and test sets. So it's not easy to separate them <a href=\"/zero92\">@zero92</a>. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "537794": "https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming",
    "540632": "I've modified the code to show the feature importances. Sure enough, various \"mean\" features are the most \"dangerous\".\n[![features](https://www.kaggleusercontent.com/kf/14989740/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..TRRcw0v2keYZBz0dkDvIlw.knDW1Jy3uRYKZdrqtyG_NTGbxKsGeJ8jhKW69_pXSWvj1Wsk3gxdimuHNyLjUwXh-JRI6n4VwWgNn6uIuEr4Daoqee13AD64BYoARHW2sMHXObhI5g27rgTBMwKD3sHmD49zwsamsUyakah84ezFHIKqKDTxGvuTwDgv6jO9Ti_8iRt1b_v-8mxqDz5Dk8xl.9sLfPTcqzxxcKn626_Bc0w/__results___files/__results___10_0.png)](https://www.kaggle.com/tunguz/lanl-adversarial-validation-shakeup-is-coming)\n",
    "537943": "My adversarial validation AUC is around 0.85 as well :)",
    "537883": "As others have noted, there is a significant difference in the acoustic mean between the two datasets, and I believe AV is reliant on this and other mean-weighted features. Personally, I think the best results in this competition will not rely on these features, and train/test will consequently be more similar. I could be wrong, but those are my thoughts!",
    "537867": "You are as addicted to AV as I am to GP ;)",
    "538965": "I think your high AUC may just be due to curse of dimensionality. I could be wrong. How many features does it have again?",
    "537884": ""
  }
}