{
  "id": 346145,
  "title": "[Rant] Why were we ever given the test data?",
  "url": "/competitions/amex-default-prediction/discussion/346145",
  "author_name": "Carl McBride Ellis",
  "post_date": "2022-08-18T06:04:15.514000",
  "votes": 49,
  "comment_count": 18,
  "views": 0,
  "content": "<p>As this competition draws to a close I find myself reflecting on the question of why in this (and the vast majority of other kaggle competitions) are we ever provided with the test data? The whole point of ML is to create a model that performs well on <strong>unseen</strong> future data, and the point of a competition is to put those models to the test on such data. When a competition is launched the first thing we do is try to work out the Public/Private partitioning of the test data, and then perform KS tests and/or adversarial validation of these partitions and the features therein with respect to the training data. However, this eventually leads to the most pernicious form of bad modeling, namely data leakage. A model that has been crafted by in any way looking ahead at the test data it will be receiving may very well perform better on kaggle, but it is in its essence a bad model, especially from the point of view of production. Indeed the problem of train-test contamination is even highlighted in the <a href=\"https://www.kaggle.com/code/alexisbcook/data-leakage/tutorial\" target=\"_blank\">kaggle tutorial on data leakage</a>.</p>\n<p>Having the test data also places undue emphasis on ensembling. Programmatic ensembling (for example stacking generalization using OOF predictions on the training data, or blending the hold-out predictions that did not form part of the CV) is a wonderful finishing touch. However, having direct access to the test data leads to a deluge of trivial notebooks that simply harvest submission files which are then 'blended' via trial-and-error linear combinations into high scoring submissions that make a mockery of both the Public (and quite potentially the Private) leaderboards, as well as the Notebooks section (such notebooks even get Silver and Gold medals!)</p>\n<p>I am quite happy to be enlightened, but at the moment I can see no redeeming reason for ever having prior access to the test data.</p>\n<p>All the best,<br>\ncarl</p>\n<p>PS: As I mention below, the solution is almost trivial; kaggle could simply provide a very small <code>test.csv</code> sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true <code>test.csv</code>.</p>",
  "messages": [
    {
      "id": 1904344,
      "postDate": "2022-08-18T06:04:15.513Z",
      "content": "<p>As this competition draws to a close I find myself reflecting on the question of why in this (and the vast majority of other kaggle competitions) are we ever provided with the test data? The whole point of ML is to create a model that performs well on <strong>unseen</strong> future data, and the point of a competition is to put those models to the test on such data. When a competition is launched the first thing we do is try to work out the Public/Private partitioning of the test data, and then perform KS tests and/or adversarial validation of these partitions and the features therein with respect to the training data. However, this eventually leads to the most pernicious form of bad modeling, namely data leakage. A model that has been crafted by in any way looking ahead at the test data it will be receiving may very well perform better on kaggle, but it is in its essence a bad model, especially from the point of view of production. Indeed the problem of train-test contamination is even highlighted in the <a href=\"https://www.kaggle.com/code/alexisbcook/data-leakage/tutorial\" target=\"_blank\">kaggle tutorial on data leakage</a>.</p>\n<p>Having the test data also places undue emphasis on ensembling. Programmatic ensembling (for example stacking generalization using OOF predictions on the training data, or blending the hold-out predictions that did not form part of the CV) is a wonderful finishing touch. However, having direct access to the test data leads to a deluge of trivial notebooks that simply harvest submission files which are then 'blended' via trial-and-error linear combinations into high scoring submissions that make a mockery of both the Public (and quite potentially the Private) leaderboards, as well as the Notebooks section (such notebooks even get Silver and Gold medals!)</p>\n<p>I am quite happy to be enlightened, but at the moment I can see no redeeming reason for ever having prior access to the test data.</p>\n<p>All the best,<br>\ncarl</p>\n<p>PS: As I mention below, the solution is almost trivial; kaggle could simply provide a very small <code>test.csv</code> sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true <code>test.csv</code>.</p>",
      "rawMarkdown": "As this competition draws to a close I find myself reflecting on the question of why in this (and the vast majority of other kaggle competitions) are we ever provided with the test data? The whole point of ML is to create a model that performs well on **unseen** future data, and the point of a competition is to put those models to the test on such data. When a competition is launched the first thing we do is try to work out the Public/Private partitioning of the test data, and then perform KS tests and/or adversarial validation of these partitions and the features therein with respect to the training data. However, this eventually leads to the most pernicious form of bad modeling, namely data leakage. A model that has been crafted by in any way looking ahead at the test data it will be receiving may very well perform better on kaggle, but it is in its essence a bad model, especially from the point of view of production. Indeed the problem of train-test contamination is even highlighted in the [kaggle tutorial on data leakage](https://www.kaggle.com/code/alexisbcook/data-leakage/tutorial).\n\nHaving the test data also places undue emphasis on ensembling. Programmatic ensembling (for example stacking generalization using OOF predictions on the training data, or blending the hold-out predictions that did not form part of the CV) is a wonderful finishing touch. However, having direct access to the test data leads to a deluge of trivial notebooks that simply harvest submission files which are then 'blended' via trial-and-error linear combinations into high scoring submissions that make a mockery of both the Public (and quite potentially the Private) leaderboards, as well as the Notebooks section (such notebooks even get Silver and Gold medals!)\n\nI am quite happy to be enlightened, but at the moment I can see no redeeming reason for ever having prior access to the test data.\n\nAll the best,\ncarl\n\nPS: As I mention below, the solution is almost trivial; kaggle could simply provide a very small `test.csv` sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true `test.csv`.",
      "votes": 49
    },
    {
      "id": 1904366,
      "postDate": "2022-08-18T06:26:21.420Z",
      "content": "<p>I am kind of neutral towards this. I guess AMEX already has some crazy model ensembles in their credit scoring system (they did have kaggle GMs working for them in top DS positions) - and this is what they might want - to enhance their current approach with some new unique ideas coming from top solutions. I am pretty sure that winning models will not be used anywhere. Having that in mind, revealing out-of-time based test set does not hurt at all.</p>",
      "rawMarkdown": "I am kind of neutral towards this. I guess AMEX already has some crazy model ensembles in their credit scoring system (they did have kaggle GMs working for them in top DS positions) - and this is what they might want - to enhance their current approach with some new unique ideas coming from top solutions. I am pretty sure that winning models will not be used anywhere. Having that in mind, revealing out-of-time based test set does not hurt at all.",
      "votes": 11,
      "replies": [
        {
          "id": 1904486,
          "postDate": "2022-08-18T08:08:55.673Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>By giving us the test data they may well end up with a whole new set of crazy model ensembles.  AmEx (or any other company for that matter) really shouldn't have any craziness at all in their credit scoring system (or any other system for that matter).</p>\n<p>All the best, and many thanks for the dataset you prepared at the start of the competition and all your insights along the way!<br>\ncarl</p>",
          "rawMarkdown": "Dear @raddar \n\nBy giving us the test data they may well end up with a whole new set of crazy model ensembles.  AmEx (or any other company for that matter) really shouldn't have any craziness at all in their credit scoring system (or any other system for that matter).\n\nAll the best, and many thanks for the dataset you prepared at the start of the competition and all your insights along the way!\ncarl",
          "votes": 1
        }
      ]
    },
    {
      "id": 1904703,
      "postDate": "2022-08-18T12:23:46.160Z",
      "content": "<p>I want to emphasize additional point - since label is delayed by 18 months, they have vast of data which are can be used as features but still has no labels available. That data can be used to improve the pipeline by many kinds of models - semi/self-supervised pretraining, data drift detection etc. Those approaches may leverage the quality of the model.</p>\n<p>As an researcher in the company which works with similar types of the data, I see a great value in this concepts.</p>",
      "rawMarkdown": "I want to emphasize additional point - since label is delayed by 18 months, they have vast of data which are can be used as features but still has no labels available. That data can be used to improve the pipeline by many kinds of models - semi/self-supervised pretraining, data drift detection etc. Those approaches may leverage the quality of the model.\n\nAs an researcher in the company which works with similar types of the data, I see a great value in this concepts.",
      "votes": 3,
      "replies": [
        {
          "id": 1904737,
          "postDate": "2022-08-18T12:48:25.460Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/pavelvod\" target=\"_blank\">@pavelvod</a> </p>\n<p>Indeed. My point is simply that kaggle should hide the test data so as to obtain a true gauge of model performance. <br>\nIt would be exceedingly easy to do; simply provide a small <code>test.csv</code> sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true <code>test.csv</code>.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @pavelvod \n\nIndeed. My point is simply that kaggle should hide the test data so as to obtain a true gauge of model performance. \nIt would be exceedingly easy to do; simply provide a small `test.csv` sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true `test.csv`.\n\nAll the best,\ncarl",
          "votes": 2
        }
      ]
    },
    {
      "id": 1909365,
      "postDate": "2022-08-22T15:08:16.450Z",
      "content": "<p>About half of our competitions are hosted in the way you suggest (Code competitions).</p>\n<p>There are two main reasons we don't host all of the competitions this way:</p>\n<ol>\n<li>In general, hosts don't use the trained models that come out of Kaggle competitions. Rather, they incorporate the techniques, tricks, etc., that the winners used to improve their productions models. In many instances, providing the test set doesn't detract from.</li>\n<li>There are many Kagglers who prefer not to participate in Code competitions, and Code competitions have limitations that may not always be advantageous. </li>\n</ol>\n<p>With these things in mind, we try to strike the right balance when a competition is Code-only or the traditional format.</p>",
      "rawMarkdown": "About half of our competitions are hosted in the way you suggest (Code competitions).\n\nThere are two main reasons we don't host all of the competitions this way:\n1. In general, hosts don't use the trained models that come out of Kaggle competitions. Rather, they incorporate the techniques, tricks, etc., that the winners used to improve their productions models. In many instances, providing the test set doesn't detract from.\n2. There are many Kagglers who prefer not to participate in Code competitions, and Code competitions have limitations that may not always be advantageous. \n\nWith these things in mind, we try to strike the right balance when a competition is Code-only or the traditional format.\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 1909502,
          "postDate": "2022-08-22T16:42:55.513Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>Thank you for your reply. </p>\n<p>Regarding the second point,  I would have thought that and API for non-forecasting competitions would be (comparatively) simple, and from the point of view of the competitors wouldn't require much coding above and that which is already in place. That said it would put a stop to the habitual torrent of trivial blends (that use no new tricks or novel techniques) that clutter the Notebooks section, and which receive undue attention to the detriment of those who publish genuine content. It would also put a stop the <em>forkmitting</em> (people who do nothing other than fork and then directly submit) of the <em>'Highest Scoring Notebook of the Day'</em> (the blends that are published invariably score slightly better on the Public LB than the original work almost by definition, otherwise they would not be published) which unduly distorts the leaderboard.</p>\n<p>From a broader perspective, not providing the test data would perhaps lead to more robust solutions with better generalization, and could eventually go some way to stemming the <a href=\"https://arxiv.org/pdf/2207.07048.pdf\" target=\"_blank\">increasing reproducibility crisis seen in machine learning</a>; it would not be the first time that techniques developed on kaggle find their  way into the ML community as a whole. Indeed only last month there was an article in Nature with the ominous subtitle that suggested that <a href=\"https://www.nature.com/articles/d41586-022-02035-w\" target=\"_blank\"> ‘data leakage’ threatens the reliability of machine-learning use across disciplines</a>.</p>\n<p>Despite the marginally higher entry barrier, maybe it would be a good thing that some time in the future all kaggle competitions become code competitions?</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @inversion \n\nThank you for your reply. \n\nRegarding the second point,  I would have thought that and API for non-forecasting competitions would be (comparatively) simple, and from the point of view of the competitors wouldn't require much coding above and that which is already in place. That said it would put a stop to the habitual torrent of trivial blends (that use no new tricks or novel techniques) that clutter the Notebooks section, and which receive undue attention to the detriment of those who publish genuine content. It would also put a stop the *forkmitting* (people who do nothing other than fork and then directly submit) of the *'Highest Scoring Notebook of the Day'* (the blends that are published invariably score slightly better on the Public LB than the original work almost by definition, otherwise they would not be published) which unduly distorts the leaderboard.\n\nFrom a broader perspective, not providing the test data would perhaps lead to more robust solutions with better generalization, and could eventually go some way to stemming the [increasing reproducibility crisis seen in machine learning](https://arxiv.org/pdf/2207.07048.pdf); it would not be the first time that techniques developed on kaggle find their  way into the ML community as a whole. Indeed only last month there was an article in Nature with the ominous subtitle that suggested that [ ‘data leakage’ threatens the reliability of machine-learning use across disciplines](https://www.nature.com/articles/d41586-022-02035-w).\n\nDespite the marginally higher entry barrier, maybe it would be a good thing that some time in the future all kaggle competitions become code competitions?\n\nAll the best,\ncarl",
          "votes": 1
        },
        {
          "id": 1909613,
          "postDate": "2022-08-22T18:45:49.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>\n<blockquote>\n  <p>(the blends that are published invariably score slightly better on the Public LB than the original work almost by definition, otherwise they would not be published)</p>\n</blockquote>\n<p>Carl and Kaggle, there is a curious public notebook <code>sort by score</code> bug that encourages Kagglers to fork notebooks and submit without changes.</p>\n<p>i think whenever someone forks a notebook and submits the identical code with identical submission.csv and identical LB score, i think the new forked notebook will appear better (i.e. ordered first) in the public notebook display when <code>sorted by score</code>.</p>\n<p>This encourages Kagglers to fork notebooks without changes because the forked notebook gets more attention than original. Perhaps when two notebooks have the same LB score, the older notebook should be displayed first when <code>sorting by score</code>.</p>\n<p>I have also noticed that people fork the highest scoring notebook, make changes that decrease the LB score but has the same visible 3 decimal digits. Then the lower scoring notebook ranks better when <code>sorting by score</code> (because sorting only uses visible digits then sorts by recent) and Kagglers' are misled into thinking the changes improved the notebook when in reality the changes decreased the model performance.</p>",
          "rawMarkdown": "@carlmcbrideellis \n>(the blends that are published invariably score slightly better on the Public LB than the original work almost by definition, otherwise they would not be published)\n\nCarl and Kaggle, there is a curious public notebook `sort by score` bug that encourages Kagglers to fork notebooks and submit without changes.\n\ni think whenever someone forks a notebook and submits the identical code with identical submission.csv and identical LB score, i think the new forked notebook will appear better (i.e. ordered first) in the public notebook display when `sorted by score`.\n\nThis encourages Kagglers to fork notebooks without changes because the forked notebook gets more attention than original. Perhaps when two notebooks have the same LB score, the older notebook should be displayed first when `sorting by score`.\n\nI have also noticed that people fork the highest scoring notebook, make changes that decrease the LB score but has the same visible 3 decimal digits. Then the lower scoring notebook ranks better when `sorting by score` (because sorting only uses visible digits then sorts by recent) and Kagglers' are misled into thinking the changes improved the notebook when in reality the changes decreased the model performance.",
          "votes": 4
        },
        {
          "id": 1909637,
          "postDate": "2022-08-22T19:26:28.113Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>\n<p>Regarding the \"identical score\" bug, I think that was the (indirect) message of the notebook <a href=\"https://www.kaggle.com/code/scumufeng/notebook-score-sorting-test-display-bug\" target=\"_blank\">\"Notebook Score Sorting Test--Display bug\"</a> after seeing the publication of the bizzarley effective <a href=\"https://www.kaggle.com/code/cbeaud/american-express-basic-scaling-submission\" target=\"_blank\">\"American Express ✨ Basic Scaling Submission\"</a> notebook.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @cdeotte\n  \nRegarding the \"identical score\" bug, I think that was the (indirect) message of the notebook [\"Notebook Score Sorting Test--Display bug\"](https://www.kaggle.com/code/scumufeng/notebook-score-sorting-test-display-bug) after seeing the publication of the bizzarley effective [\"American Express ✨ Basic Scaling Submission\"](https://www.kaggle.com/code/cbeaud/american-express-basic-scaling-submission) notebook.\n\nAll the best,\ncarl",
          "votes": 1
        }
      ]
    },
    {
      "id": 1904610,
      "postDate": "2022-08-18T10:45:21.823Z",
      "content": "<p>Well, I think it is not always necessary for a solution to work on <strong>unseen</strong> data. If AMEX wants to make real predictions for their current set of customers, they also see the test data and could use it in their prediction pipeline, although this probably makes things slightly more complicated.</p>",
      "rawMarkdown": "Well, I think it is not always necessary for a solution to work on **unseen** data. If AMEX wants to make real predictions for their current set of customers, they also see the test data and could use it in their prediction pipeline, although this probably makes things slightly more complicated.",
      "votes": 1,
      "replies": [
        {
          "id": 1904683,
          "postDate": "2022-08-18T11:56:57.560Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> </p>\n<p>I would <em>never</em> trust the predictions of a model that was trained (in any way) on the very data it will be asked to predict</p>\n<p>All the best,<br>\ncarl </p>",
          "rawMarkdown": "Dear @fritzcremer \n\nI would *never* trust the predictions of a model that was trained (in any way) on the very data it will be asked to predict\n\nAll the best,\ncarl "
        },
        {
          "id": 1904701,
          "postDate": "2022-08-18T12:19:23.460Z",
          "content": "<p>Hmm but why not, if the training involves e.g. unsupervised pre-training?</p>",
          "rawMarkdown": "Hmm but why not, if the training involves e.g. unsupervised pre-training?",
          "votes": 2
        },
        {
          "id": 1904706,
          "postDate": "2022-08-18T12:27:36.827Z",
          "content": "<p>Because the results will be overly optimistic, and the moment the model is rolled-out on truly unseen data its performance will drop (although not by as much as it would if there were target leakage as well). There is noting at all wrong with unsupervised pre-training, as long as it is not done using the actual test data!</p>",
          "rawMarkdown": "Because the results will be overly optimistic, and the moment the model is rolled-out on truly unseen data its performance will drop (although not by as much as it would if there were target leakage as well). There is noting at all wrong with unsupervised pre-training, as long as it is not done using the actual test data!"
        },
        {
          "id": 1905044,
          "postDate": "2022-08-18T17:44:47.260Z",
          "content": "<p>Yes, I see what you mean. I think this is an issue with e.g. the problem provided by AMEX, since they can't rerun their whole training pipeline, before making a prediction for a new user. However, when Inference happens rarely and on large scale, each new inference set could be included in the training pipeline, thus making conditions equal for validation and test data.</p>",
          "rawMarkdown": "Yes, I see what you mean. I think this is an issue with e.g. the problem provided by AMEX, since they can't rerun their whole training pipeline, before making a prediction for a new user. However, when Inference happens rarely and on large scale, each new inference set could be included in the training pipeline, thus making conditions equal for validation and test data."
        }
      ]
    },
    {
      "id": 1904733,
      "postDate": "2022-08-18T12:46:53.117Z",
      "content": "<p>I'll go against the gradient and say maybe they don't care too much about the winning model, and maybe its just for marketing sake because Kaggle competitions like these tend to be some of the most popular, and they are advertising they have roles open for hiring so want more applicants to see it.</p>\n<p>There's ~4600 teams here, compared to all other active competitions which are average around ~1000..</p>",
      "rawMarkdown": "I'll go against the gradient and say maybe they don't care too much about the winning model, and maybe its just for marketing sake because Kaggle competitions like these tend to be some of the most popular, and they are advertising they have roles open for hiring so want more applicants to see it.\n\nThere's ~4600 teams here, compared to all other active competitions which are average around ~1000..",
      "replies": [
        {
          "id": 1904743,
          "postDate": "2022-08-18T12:53:45.957Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> </p>\n<p>I am oft reminded of  <a href=\"https://www.mit.edu/~xela/tao.html\" target=\"_blank\">\"<em>The Tao of Programming</em>\"</a> in particular section 3.1; the most valuable things are new ideas…</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @julianmukaj \n\nI am oft reminded of  [\"*The Tao of Programming*\"](https://www.mit.edu/~xela/tao.html) in particular section 3.1; the most valuable things are new ideas...\n\nAll the best,\ncarl",
          "votes": 2
        }
      ]
    },
    {
      "id": 1909017,
      "postDate": "2022-08-22T08:14:02.247Z",
      "content": "<p>My humble opinion is, “Data solution does not always work as planned”.  AMEX may use historical data to make informed decisions and align their strategic approach.</p>",
      "rawMarkdown": "My humble opinion is, “Data solution does not always work as planned”.  AMEX may use historical data to make informed decisions and align their strategic approach."
    },
    {
      "id": 1910389,
      "postDate": "2022-08-23T11:36:11.917Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1910736,
          "postDate": "2022-08-23T16:34:15.320Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a> </p>\n<p>To dispel any confusion; I am bemoaning the availability of the <code>test.csv</code> file, and no other aspect of competitions.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @gehallak \n\nTo dispel any confusion; I am bemoaning the availability of the `test.csv` file, and no other aspect of competitions.\n\nAll the best,\ncarl"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1904366,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-08-18T06:26:21.420000",
      "content": "<p>I am kind of neutral towards this. I guess AMEX already has some crazy model ensembles in their credit scoring system (they did have kaggle GMs working for them in top DS positions) - and this is what they might want - to enhance their current approach with some new unique ideas coming from top solutions. I am pretty sure that winning models will not be used anywhere. Having that in mind, revealing out-of-time based test set does not hurt at all.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1904486,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-18T08:08:55.673000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>By giving us the test data they may well end up with a whole new set of crazy model ensembles.  AmEx (or any other company for that matter) really shouldn't have any craziness at all in their credit scoring system (or any other system for that matter).</p>\n<p>All the best, and many thanks for the dataset you prepared at the start of the competition and all your insights along the way!<br>\ncarl</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1904703,
      "author_name": "Pavel Vodolazov",
      "author_url": "",
      "post_date": "2022-08-18T12:23:46.160000",
      "content": "<p>I want to emphasize additional point - since label is delayed by 18 months, they have vast of data which are can be used as features but still has no labels available. That data can be used to improve the pipeline by many kinds of models - semi/self-supervised pretraining, data drift detection etc. Those approaches may leverage the quality of the model.</p>\n<p>As an researcher in the company which works with similar types of the data, I see a great value in this concepts.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1904737,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-18T12:48:25.460000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/pavelvod\" target=\"_blank\">@pavelvod</a> </p>\n<p>Indeed. My point is simply that kaggle should hide the test data so as to obtain a true gauge of model performance. <br>\nIt would be exceedingly easy to do; simply provide a small <code>test.csv</code> sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true <code>test.csv</code>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1909365,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2022-08-22T15:08:16.450000",
      "content": "<p>About half of our competitions are hosted in the way you suggest (Code competitions).</p>\n<p>There are two main reasons we don't host all of the competitions this way:</p>\n<ol>\n<li>In general, hosts don't use the trained models that come out of Kaggle competitions. Rather, they incorporate the techniques, tricks, etc., that the winners used to improve their productions models. In many instances, providing the test set doesn't detract from.</li>\n<li>There are many Kagglers who prefer not to participate in Code competitions, and Code competitions have limitations that may not always be advantageous. </li>\n</ol>\n<p>With these things in mind, we try to strike the right balance when a competition is Code-only or the traditional format.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1909502,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-22T16:42:55.513000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>Thank you for your reply. </p>\n<p>Regarding the second point,  I would have thought that and API for non-forecasting competitions would be (comparatively) simple, and from the point of view of the competitors wouldn't require much coding above and that which is already in place. That said it would put a stop to the habitual torrent of trivial blends (that use no new tricks or novel techniques) that clutter the Notebooks section, and which receive undue attention to the detriment of those who publish genuine content. It would also put a stop the <em>forkmitting</em> (people who do nothing other than fork and then directly submit) of the <em>'Highest Scoring Notebook of the Day'</em> (the blends that are published invariably score slightly better on the Public LB than the original work almost by definition, otherwise they would not be published) which unduly distorts the leaderboard.</p>\n<p>From a broader perspective, not providing the test data would perhaps lead to more robust solutions with better generalization, and could eventually go some way to stemming the <a href=\"https://arxiv.org/pdf/2207.07048.pdf\" target=\"_blank\">increasing reproducibility crisis seen in machine learning</a>; it would not be the first time that techniques developed on kaggle find their  way into the ML community as a whole. Indeed only last month there was an article in Nature with the ominous subtitle that suggested that <a href=\"https://www.nature.com/articles/d41586-022-02035-w\" target=\"_blank\"> ‘data leakage’ threatens the reliability of machine-learning use across disciplines</a>.</p>\n<p>Despite the marginally higher entry barrier, maybe it would be a good thing that some time in the future all kaggle competitions become code competitions?</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1909613,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-22T18:45:49.137000",
          "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>\n<blockquote>\n  <p>(the blends that are published invariably score slightly better on the Public LB than the original work almost by definition, otherwise they would not be published)</p>\n</blockquote>\n<p>Carl and Kaggle, there is a curious public notebook <code>sort by score</code> bug that encourages Kagglers to fork notebooks and submit without changes.</p>\n<p>i think whenever someone forks a notebook and submits the identical code with identical submission.csv and identical LB score, i think the new forked notebook will appear better (i.e. ordered first) in the public notebook display when <code>sorted by score</code>.</p>\n<p>This encourages Kagglers to fork notebooks without changes because the forked notebook gets more attention than original. Perhaps when two notebooks have the same LB score, the older notebook should be displayed first when <code>sorting by score</code>.</p>\n<p>I have also noticed that people fork the highest scoring notebook, make changes that decrease the LB score but has the same visible 3 decimal digits. Then the lower scoring notebook ranks better when <code>sorting by score</code> (because sorting only uses visible digits then sorts by recent) and Kagglers' are misled into thinking the changes improved the notebook when in reality the changes decreased the model performance.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1909637,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-22T19:26:28.113000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>\n<p>Regarding the \"identical score\" bug, I think that was the (indirect) message of the notebook <a href=\"https://www.kaggle.com/code/scumufeng/notebook-score-sorting-test-display-bug\" target=\"_blank\">\"Notebook Score Sorting Test--Display bug\"</a> after seeing the publication of the bizzarley effective <a href=\"https://www.kaggle.com/code/cbeaud/american-express-basic-scaling-submission\" target=\"_blank\">\"American Express ✨ Basic Scaling Submission\"</a> notebook.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1904610,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2022-08-18T10:45:21.823000",
      "content": "<p>Well, I think it is not always necessary for a solution to work on <strong>unseen</strong> data. If AMEX wants to make real predictions for their current set of customers, they also see the test data and could use it in their prediction pipeline, although this probably makes things slightly more complicated.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1904683,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-18T11:56:57.560000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> </p>\n<p>I would <em>never</em> trust the predictions of a model that was trained (in any way) on the very data it will be asked to predict</p>\n<p>All the best,<br>\ncarl </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1904701,
          "author_name": "Fritz Cremer",
          "author_url": "",
          "post_date": "2022-08-18T12:19:23.460000",
          "content": "<p>Hmm but why not, if the training involves e.g. unsupervised pre-training?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1904706,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-18T12:27:36.827000",
          "content": "<p>Because the results will be overly optimistic, and the moment the model is rolled-out on truly unseen data its performance will drop (although not by as much as it would if there were target leakage as well). There is noting at all wrong with unsupervised pre-training, as long as it is not done using the actual test data!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1905044,
          "author_name": "Fritz Cremer",
          "author_url": "",
          "post_date": "2022-08-18T17:44:47.260000",
          "content": "<p>Yes, I see what you mean. I think this is an issue with e.g. the problem provided by AMEX, since they can't rerun their whole training pipeline, before making a prediction for a new user. However, when Inference happens rarely and on large scale, each new inference set could be included in the training pipeline, thus making conditions equal for validation and test data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1904733,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2022-08-18T12:46:53.117000",
      "content": "<p>I'll go against the gradient and say maybe they don't care too much about the winning model, and maybe its just for marketing sake because Kaggle competitions like these tend to be some of the most popular, and they are advertising they have roles open for hiring so want more applicants to see it.</p>\n<p>There's ~4600 teams here, compared to all other active competitions which are average around ~1000..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1904743,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-18T12:53:45.957000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> </p>\n<p>I am oft reminded of  <a href=\"https://www.mit.edu/~xela/tao.html\" target=\"_blank\">\"<em>The Tao of Programming</em>\"</a> in particular section 3.1; the most valuable things are new ideas…</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1909017,
      "author_name": "Shakeel Bhatti",
      "author_url": "",
      "post_date": "2022-08-22T08:14:02.247000",
      "content": "<p>My humble opinion is, “Data solution does not always work as planned”.  AMEX may use historical data to make informed decisions and align their strategic approach.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1910389,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-23T11:36:11.917000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1910736,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-23T16:34:15.320000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a> </p>\n<p>To dispel any confusion; I am bemoaning the availability of the <code>test.csv</code> file, and no other aspect of competitions.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1904344": "As this competition draws to a close I find myself reflecting on the question of why in this (and the vast majority of other kaggle competitions) are we ever provided with the test data? The whole point of ML is to create a model that performs well on **unseen** future data, and the point of a competition is to put those models to the test on such data. When a competition is launched the first thing we do is try to work out the Public/Private partitioning of the test data, and then perform KS tests and/or adversarial validation of these partitions and the features therein with respect to the training data. However, this eventually leads to the most pernicious form of bad modeling, namely data leakage. A model that has been crafted by in any way looking ahead at the test data it will be receiving may very well perform better on kaggle, but it is in its essence a bad model, especially from the point of view of production. Indeed the problem of train-test contamination is even highlighted in the [kaggle tutorial on data leakage](https://www.kaggle.com/code/alexisbcook/data-leakage/tutorial).\n\nHaving the test data also places undue emphasis on ensembling. Programmatic ensembling (for example stacking generalization using OOF predictions on the training data, or blending the hold-out predictions that did not form part of the CV) is a wonderful finishing touch. However, having direct access to the test data leads to a deluge of trivial notebooks that simply harvest submission files which are then 'blended' via trial-and-error linear combinations into high scoring submissions that make a mockery of both the Public (and quite potentially the Private) leaderboards, as well as the Notebooks section (such notebooks even get Silver and Gold medals!)\n\nI am quite happy to be enlightened, but at the moment I can see no redeeming reason for ever having prior access to the test data.\n\nAll the best,\ncarl\n\nPS: As I mention below, the solution is almost trivial; kaggle could simply provide a very small `test.csv` sufficient to ensure the correct working of a script, and upon submission they would swap that file out for the true `test.csv`.",
    "1904366": "I am kind of neutral towards this. I guess AMEX already has some crazy model ensembles in their credit scoring system (they did have kaggle GMs working for them in top DS positions) - and this is what they might want - to enhance their current approach with some new unique ideas coming from top solutions. I am pretty sure that winning models will not be used anywhere. Having that in mind, revealing out-of-time based test set does not hurt at all.",
    "1904703": "I want to emphasize additional point - since label is delayed by 18 months, they have vast of data which are can be used as features but still has no labels available. That data can be used to improve the pipeline by many kinds of models - semi/self-supervised pretraining, data drift detection etc. Those approaches may leverage the quality of the model.\n\nAs an researcher in the company which works with similar types of the data, I see a great value in this concepts.",
    "1909365": "About half of our competitions are hosted in the way you suggest (Code competitions).\n\nThere are two main reasons we don't host all of the competitions this way:\n1. In general, hosts don't use the trained models that come out of Kaggle competitions. Rather, they incorporate the techniques, tricks, etc., that the winners used to improve their productions models. In many instances, providing the test set doesn't detract from.\n2. There are many Kagglers who prefer not to participate in Code competitions, and Code competitions have limitations that may not always be advantageous. \n\nWith these things in mind, we try to strike the right balance when a competition is Code-only or the traditional format.\n\n",
    "1904610": "Well, I think it is not always necessary for a solution to work on **unseen** data. If AMEX wants to make real predictions for their current set of customers, they also see the test data and could use it in their prediction pipeline, although this probably makes things slightly more complicated.",
    "1904733": "I'll go against the gradient and say maybe they don't care too much about the winning model, and maybe its just for marketing sake because Kaggle competitions like these tend to be some of the most popular, and they are advertising they have roles open for hiring so want more applicants to see it.\n\nThere's ~4600 teams here, compared to all other active competitions which are average around ~1000..",
    "1909017": "My humble opinion is, “Data solution does not always work as planned”.  AMEX may use historical data to make informed decisions and align their strategic approach.",
    "1910389": ""
  }
}