{
  "id": 254441,
  "title": "Train vs test: drastic difference",
  "url": "/competitions/seti-breakthrough-listen/discussion/254441",
  "author_name": "gooron",
  "post_date": "2021-07-21T20:38:01.034000",
  "votes": 44,
  "comment_count": 16,
  "views": 0,
  "content": "<p>The competition was just restarted and almost immediately it became clear that <br>\n1) it is much more difficult: CV ~ 0.98… becomes ~ 0.85…; <br>\n2) there is a rather large difference between CV and LB ~ 0.1-0.15. <br>\n  There is a suspicion that (2) may be the result of a significant difference between the training and test data. <br>\n To test this, I conducted a simple standard test, namely, I tried to train a binary classifier that distinguishes between training and test data using the same pipe-line as for the main task.<br>\n Here are some details (I believe, not very significant):<br>\n  np.hstack([p[0,:,:], p[2,:,:], p[4,:,:]]) is used as an input and pre-trained efficientnetv2 is used as a model; no mixing, BCEWithLogitsLoss as loss function and AdamW as optimizer.<br>\nHere are the main results of my experiments:<br>\n (1) The neural network easily learns to distinguish between training and test data, after just a few iterations, <br>\n an accuracy of 0.98… and roc_auc &gt; 0.99 is achieved.   <br>\n (2) This property does not depend significantly on the size of the images (i checked several: (273,256), … (68,64))   <br>\n    or on normalization (a la CPMP proposed). Moreover, the same results are observed if we take only the calibration part<br>\n    of the cadence as an input: np.hstack([p[1,:,:], p[3,:,:], p[5,:,:]])<br>\n   It is worth noting that the leaky old data showed standard results for a well-posed task: <br>\n         accuracy ~ 0.58 = (1 - ntest/ntrain)  and roc_auc ~ 0.5<br>\n  There is a question to the organizers, is this \"feature\" planned or is the next restart waiting for us?<br>\n  If the first is true, then this is formally one of the most difficult competitions.     </p>",
  "messages": [
    {
      "id": 1396150,
      "postDate": "2021-07-21T20:38:01.033Z",
      "content": "<p>The competition was just restarted and almost immediately it became clear that <br>\n1) it is much more difficult: CV ~ 0.98… becomes ~ 0.85…; <br>\n2) there is a rather large difference between CV and LB ~ 0.1-0.15. <br>\n  There is a suspicion that (2) may be the result of a significant difference between the training and test data. <br>\n To test this, I conducted a simple standard test, namely, I tried to train a binary classifier that distinguishes between training and test data using the same pipe-line as for the main task.<br>\n Here are some details (I believe, not very significant):<br>\n  np.hstack([p[0,:,:], p[2,:,:], p[4,:,:]]) is used as an input and pre-trained efficientnetv2 is used as a model; no mixing, BCEWithLogitsLoss as loss function and AdamW as optimizer.<br>\nHere are the main results of my experiments:<br>\n (1) The neural network easily learns to distinguish between training and test data, after just a few iterations, <br>\n an accuracy of 0.98… and roc_auc &gt; 0.99 is achieved.   <br>\n (2) This property does not depend significantly on the size of the images (i checked several: (273,256), … (68,64))   <br>\n    or on normalization (a la CPMP proposed). Moreover, the same results are observed if we take only the calibration part<br>\n    of the cadence as an input: np.hstack([p[1,:,:], p[3,:,:], p[5,:,:]])<br>\n   It is worth noting that the leaky old data showed standard results for a well-posed task: <br>\n         accuracy ~ 0.58 = (1 - ntest/ntrain)  and roc_auc ~ 0.5<br>\n  There is a question to the organizers, is this \"feature\" planned or is the next restart waiting for us?<br>\n  If the first is true, then this is formally one of the most difficult competitions.     </p>",
      "rawMarkdown": "   The competition was just restarted and almost immediately it became clear that \n\n  1) it is much more difficult: CV ~ 0.98... becomes ~ 0.85...; \n  2) there is a rather large difference between CV and LB ~ 0.1-0.15. \n\n    There is a suspicion that (2) may be the result of a significant difference between the training and test data. \n    To test this, I conducted a simple standard test, namely, I tried to train a binary classifier that distinguishes between training and test data using the same pipe-line as for the main task.\n    Here are some details (I believe, not very significant):\n \n   np.hstack([p[0,:,:], p[2,:,:], p[4,:,:]]) is used as an input and pre-trained efficientnetv2 is used as a model; no mixing, BCEWithLogitsLoss as loss function and AdamW as optimizer.\n\nHere are the main results of my experiments:\n\n   (1) The neural network easily learns to distinguish between training and test data, after just a few iterations, \n    an accuracy of 0.98... and roc_auc > 0.99 is achieved.   \n\n   (2) This property does not depend significantly on the size of the images (i checked several: (273,256), ... (68,64))   \n       or on normalization (a la CPMP proposed). Moreover, the same results are observed if we take only the calibration part\n       of the cadence as an input: np.hstack([p[1,:,:], p[3,:,:], p[5,:,:]])\n \n    It is worth noting that the leaky old data showed standard results for a well-posed task: \n            accuracy ~ 0.58 = (1 - ntest/ntrain)  and roc_auc ~ 0.5\n\n    There is a question to the organizers, is this \"feature\" planned or is the next restart waiting for us?\n\n    If the first is true, then this is formally one of the most difficult competitions.     \n",
      "votes": 44
    },
    {
      "id": 1401009,
      "postDate": "2021-07-26T20:08:14.747Z",
      "content": "<p>I wouldn't worry about another reset. The gap and extra difficulty are intended.</p>\n<p>We've reduced the intensities of the injected signals compared to the noise, which should be part of the reason why models are doing less well in general. We also snuck some samples into the test set that are generated differently from the ones in the training set, but they still adhere to the description in the overview page:</p>\n<blockquote>\n  <p>but what they do have in common is that they are only present in some or all of the “A” observations</p>\n</blockquote>",
      "rawMarkdown": "I wouldn't worry about another reset. The gap and extra difficulty are intended.\n\nWe've reduced the intensities of the injected signals compared to the noise, which should be part of the reason why models are doing less well in general. We also snuck some samples into the test set that are generated differently from the ones in the training set, but they still adhere to the description in the overview page:\n\n> but what they do have in common is that they are only present in some or all of the “A” observations",
      "votes": 14,
      "replies": [
        {
          "id": 1401029,
          "postDate": "2021-07-26T20:42:39.053Z",
          "content": "<p>thank you, this is what one suspected but still it's good to have a solid statement from the organizers</p>",
          "rawMarkdown": "thank you, this is what one suspected but still it's good to have a solid statement from the organizers"
        },
        {
          "id": 1401058,
          "postDate": "2021-07-26T21:38:30.003Z",
          "content": "<p>Thanks for the clarification. One further question: is there also a deliberate \"gap\" between the public and private test sets, or are they pretty much randomly separated?</p>",
          "rawMarkdown": "Thanks for the clarification. One further question: is there also a deliberate \"gap\" between the public and private test sets, or are they pretty much randomly separated?"
        },
        {
          "id": 1401345,
          "postDate": "2021-07-27T08:35:48.693Z",
          "content": "<p><code>They are only present in some or all of the “A” observations</code></p>\n<p>Maybe this is the key for this competition, I'm curious to read the top solutions on how to learn from this after competion.</p>",
          "rawMarkdown": "`They are only present in some or all of the “A” observations`\n\nMaybe this is the key for this competition, I'm curious to read the top solutions on how to learn from this after competion."
        },
        {
          "id": 1402078,
          "postDate": "2021-07-27T20:29:21.317Z",
          "content": "<p>Can you clarify what you mean by \"deliberate gap\"? The difference between the test set and training set is deliberate, but any score/metric difference between CV and LB will depend on the model.</p>",
          "rawMarkdown": "Can you clarify what you mean by \"deliberate gap\"? The difference between the test set and training set is deliberate, but any score/metric difference between CV and LB will depend on the model.",
          "votes": 1
        },
        {
          "id": 1402102,
          "postDate": "2021-07-27T21:07:05.627Z",
          "content": "<p>What I am asking about is the difference, if due to something other than random selection, between the public test data and the private test data, not the difference between the training data and the test data (public or private).</p>",
          "rawMarkdown": "What I am asking about is the difference, if due to something other than random selection, between the public test data and the private test data, not the difference between the training data and the test data (public or private)."
        },
        {
          "id": 1402155,
          "postDate": "2021-07-27T22:47:52.377Z",
          "content": "<p>Ah, the split between public and private test is completely random iirc.</p>",
          "rawMarkdown": "Ah, the split between public and private test is completely random iirc.",
          "votes": 10
        }
      ]
    },
    {
      "id": 1397229,
      "postDate": "2021-07-23T00:19:48.437Z",
      "content": "<p>Thank you for sharing interesting point. <a href=\"https://www.kaggle.com/tomooinubushi/train-vs-test-lightgbm\" target=\"_blank\">I confirmed your result with light GBM</a>. I found basic statistical features (e.g., min, max, and nunique) are different between train and test sets.<br>\nDo anybody have ideas to correct these differences?</p>",
      "rawMarkdown": "Thank you for sharing interesting point. [I confirmed your result with light GBM](https://www.kaggle.com/tomooinubushi/train-vs-test-lightgbm). I found basic statistical features (e.g., min, max, and nunique) are different between train and test sets.\nDo anybody have ideas to correct these differences?",
      "votes": 3,
      "replies": [
        {
          "id": 1397279,
          "postDate": "2021-07-23T02:54:02.953Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1413893,
          "postDate": "2021-08-02T20:52:23.987Z",
          "content": "<p>Using a simple 2-fold CV setup, I trained an ensemble of lightgbm models to predict test set membership using an assortment of image-related features extracted from the new data, and got about 0.9536 AUC score.</p>\n<p>This suggests that the train and test data do differ by much more than just a smattering of new-style alien signals sprinkled into the test set.</p>",
          "rawMarkdown": "Using a simple 2-fold CV setup, I trained an ensemble of lightgbm models to predict test set membership using an assortment of image-related features extracted from the new data, and got about 0.9536 AUC score.\n\nThis suggests that the train and test data do differ by much more than just a smattering of new-style alien signals sprinkled into the test set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1396363,
      "postDate": "2021-07-22T03:49:04.050Z",
      "content": "<p>Yes. I agree with you.<br>\nAfter I notice that, I had try domain adaptation using DANN. <br>\n(DANN's main Idea is find the embedding-space that cannot distinguish with nueral network)<br>\nBut it doesn't give me good results.</p>",
      "rawMarkdown": "Yes. I agree with you.\nAfter I notice that, I had try domain adaptation using DANN. \n(DANN's main Idea is find the embedding-space that cannot distinguish with nueral network)\nBut it doesn't give me good results.",
      "votes": 4
    },
    {
      "id": 1396310,
      "postDate": "2021-07-22T02:18:28.677Z",
      "content": "<p>I noticed it too. The competition now has a new puzzle and its name is <em>domain shift</em>.</p>",
      "rawMarkdown": "I noticed it too. The competition now has a new puzzle and its name is *domain shift*.",
      "votes": 4,
      "replies": [
        {
          "id": 1396371,
          "postDate": "2021-07-22T04:08:41.860Z",
          "content": "<p>Kaggle Competition reminds me of GitHub+VSCode -&gt; Copilot.</p>",
          "rawMarkdown": "Kaggle Competition reminds me of GitHub+VSCode -> Copilot.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1398590,
      "postDate": "2021-07-24T10:02:35.637Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/gooron\" target=\"_blank\">@gooron</a> , hope this \"feature\" you mentioned is not another leak, if it`s not leak I guess maybe to find how to fix this large domain shift is key for the new data.</p>",
      "rawMarkdown": "Thanks for sharing @gooron , hope this \"feature\" you mentioned is not another leak, if it`s not leak I guess maybe to find how to fix this large domain shift is key for the new data.",
      "votes": 1
    },
    {
      "id": 1397119,
      "postDate": "2021-07-22T19:31:04.737Z",
      "content": "<p>I agree. its good.</p>",
      "rawMarkdown": "I agree. its good."
    },
    {
      "id": 1397496,
      "postDate": "2021-07-23T08:42:41.940Z",
      "content": "<p>Thanks for Sharing this point</p>",
      "rawMarkdown": "Thanks for Sharing this point"
    }
  ],
  "comments": [
    {
      "id": 1401009,
      "author_name": "Yuhong Chen",
      "author_url": "",
      "post_date": "2021-07-26T20:08:14.747000",
      "content": "<p>I wouldn't worry about another reset. The gap and extra difficulty are intended.</p>\n<p>We've reduced the intensities of the injected signals compared to the noise, which should be part of the reason why models are doing less well in general. We also snuck some samples into the test set that are generated differently from the ones in the training set, but they still adhere to the description in the overview page:</p>\n<blockquote>\n  <p>but what they do have in common is that they are only present in some or all of the “A” observations</p>\n</blockquote>",
      "votes": 14,
      "replies": [
        {
          "id": 1401029,
          "author_name": "gooron",
          "author_url": "",
          "post_date": "2021-07-26T20:42:39.053000",
          "content": "<p>thank you, this is what one suspected but still it's good to have a solid statement from the organizers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1401058,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2021-07-26T21:38:30.003000",
          "content": "<p>Thanks for the clarification. One further question: is there also a deliberate \"gap\" between the public and private test sets, or are they pretty much randomly separated?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1401345,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-07-27T08:35:48.693000",
          "content": "<p><code>They are only present in some or all of the “A” observations</code></p>\n<p>Maybe this is the key for this competition, I'm curious to read the top solutions on how to learn from this after competion.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1402078,
          "author_name": "Yuhong Chen",
          "author_url": "",
          "post_date": "2021-07-27T20:29:21.317000",
          "content": "<p>Can you clarify what you mean by \"deliberate gap\"? The difference between the test set and training set is deliberate, but any score/metric difference between CV and LB will depend on the model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1402102,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2021-07-27T21:07:05.627000",
          "content": "<p>What I am asking about is the difference, if due to something other than random selection, between the public test data and the private test data, not the difference between the training data and the test data (public or private).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1402155,
          "author_name": "Yuhong Chen",
          "author_url": "",
          "post_date": "2021-07-27T22:47:52.377000",
          "content": "<p>Ah, the split between public and private test is completely random iirc.</p>",
          "votes": 10,
          "replies": []
        }
      ]
    },
    {
      "id": 1397229,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-07-23T00:19:48.437000",
      "content": "<p>Thank you for sharing interesting point. <a href=\"https://www.kaggle.com/tomooinubushi/train-vs-test-lightgbm\" target=\"_blank\">I confirmed your result with light GBM</a>. I found basic statistical features (e.g., min, max, and nunique) are different between train and test sets.<br>\nDo anybody have ideas to correct these differences?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1397279,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-23T02:54:02.953000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1413893,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2021-08-02T20:52:23.987000",
          "content": "<p>Using a simple 2-fold CV setup, I trained an ensemble of lightgbm models to predict test set membership using an assortment of image-related features extracted from the new data, and got about 0.9536 AUC score.</p>\n<p>This suggests that the train and test data do differ by much more than just a smattering of new-style alien signals sprinkled into the test set.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1396363,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-07-22T03:49:04.050000",
      "content": "<p>Yes. I agree with you.<br>\nAfter I notice that, I had try domain adaptation using DANN. <br>\n(DANN's main Idea is find the embedding-space that cannot distinguish with nueral network)<br>\nBut it doesn't give me good results.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1396310,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-07-22T02:18:28.677000",
      "content": "<p>I noticed it too. The competition now has a new puzzle and its name is <em>domain shift</em>.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1396371,
          "author_name": "WOOSUNG YOON",
          "author_url": "",
          "post_date": "2021-07-22T04:08:41.860000",
          "content": "<p>Kaggle Competition reminds me of GitHub+VSCode -&gt; Copilot.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1398590,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-07-24T10:02:35.637000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/gooron\" target=\"_blank\">@gooron</a> , hope this \"feature\" you mentioned is not another leak, if it`s not leak I guess maybe to find how to fix this large domain shift is key for the new data.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1397119,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-22T19:31:04.737000",
      "content": "<p>I agree. its good.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1397496,
      "author_name": "Puneet Sivananda",
      "author_url": "",
      "post_date": "2021-07-23T08:42:41.940000",
      "content": "<p>Thanks for Sharing this point</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1396150": "   The competition was just restarted and almost immediately it became clear that \n\n  1) it is much more difficult: CV ~ 0.98... becomes ~ 0.85...; \n  2) there is a rather large difference between CV and LB ~ 0.1-0.15. \n\n    There is a suspicion that (2) may be the result of a significant difference between the training and test data. \n    To test this, I conducted a simple standard test, namely, I tried to train a binary classifier that distinguishes between training and test data using the same pipe-line as for the main task.\n    Here are some details (I believe, not very significant):\n \n   np.hstack([p[0,:,:], p[2,:,:], p[4,:,:]]) is used as an input and pre-trained efficientnetv2 is used as a model; no mixing, BCEWithLogitsLoss as loss function and AdamW as optimizer.\n\nHere are the main results of my experiments:\n\n   (1) The neural network easily learns to distinguish between training and test data, after just a few iterations, \n    an accuracy of 0.98... and roc_auc > 0.99 is achieved.   \n\n   (2) This property does not depend significantly on the size of the images (i checked several: (273,256), ... (68,64))   \n       or on normalization (a la CPMP proposed). Moreover, the same results are observed if we take only the calibration part\n       of the cadence as an input: np.hstack([p[1,:,:], p[3,:,:], p[5,:,:]])\n \n    It is worth noting that the leaky old data showed standard results for a well-posed task: \n            accuracy ~ 0.58 = (1 - ntest/ntrain)  and roc_auc ~ 0.5\n\n    There is a question to the organizers, is this \"feature\" planned or is the next restart waiting for us?\n\n    If the first is true, then this is formally one of the most difficult competitions.     \n",
    "1401009": "I wouldn't worry about another reset. The gap and extra difficulty are intended.\n\nWe've reduced the intensities of the injected signals compared to the noise, which should be part of the reason why models are doing less well in general. We also snuck some samples into the test set that are generated differently from the ones in the training set, but they still adhere to the description in the overview page:\n\n> but what they do have in common is that they are only present in some or all of the “A” observations",
    "1397229": "Thank you for sharing interesting point. [I confirmed your result with light GBM](https://www.kaggle.com/tomooinubushi/train-vs-test-lightgbm). I found basic statistical features (e.g., min, max, and nunique) are different between train and test sets.\nDo anybody have ideas to correct these differences?",
    "1396363": "Yes. I agree with you.\nAfter I notice that, I had try domain adaptation using DANN. \n(DANN's main Idea is find the embedding-space that cannot distinguish with nueral network)\nBut it doesn't give me good results.",
    "1396310": "I noticed it too. The competition now has a new puzzle and its name is *domain shift*.",
    "1398590": "Thanks for sharing @gooron , hope this \"feature\" you mentioned is not another leak, if it`s not leak I guess maybe to find how to fix this large domain shift is key for the new data.",
    "1397119": "I agree. its good.",
    "1397496": "Thanks for Sharing this point"
  }
}