{
  "id": 486555,
  "title": "Phantom Run on Hidden Test Set?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/486555",
  "author_name": "",
  "post_date": "2024-03-25T13:45:01.183276500Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>When we are satisfied with a notebook, we hit Submit … we can then watch it run, see the logs, etc (kind of exciting to watch), but that run that we see is just a ruse and is only their help us to know whether the code had any issues.  And to my surprise, that run is still only using the 10 test set row dataset.</p>\n<p>There must be another hidden / phantom run that we are unable to see the details for, correct?  That is what is actually scoring our run.  </p>\n<p>This is a pretty clever way to hide the test set from us.  I was super apprehensive about exporting the test set from my \"submit\" code, thinking it was illegal to do so.  I reluctantly did a \"shape\" on it and saw it was only 10 rows.  Was relieved not to be tempted by it anymore, but was also impressed with the architecture design of the process.  </p>\n<p>This is likely why peoples code is failing a lot -- because during inference, they run out of memory.  I would suggest for folks to run inference on the train set because we were told that the test set is around the same size (90% of the size I think).</p>\n<p>(this is my first competition where I have submitted more than one solution (more than just copy/submit notebook) … this might be normal for other competitions, but it was a surprise for me… thanks!)</p>",
  "messages": [
    {
      "id": "2715428",
      "postDate": "03/25/2024 13:45:01",
      "content": "<p>When we are satisfied with a notebook, we hit Submit … we can then watch it run, see the logs, etc (kind of exciting to watch), but that run that we see is just a ruse and is only their help us to know whether the code had any issues.  And to my surprise, that run is still only using the 10 test set row dataset.</p>\n<p>There must be another hidden / phantom run that we are unable to see the details for, correct?  That is what is actually scoring our run.  </p>\n<p>This is a pretty clever way to hide the test set from us.  I was super apprehensive about exporting the test set from my \"submit\" code, thinking it was illegal to do so.  I reluctantly did a \"shape\" on it and saw it was only 10 rows.  Was relieved not to be tempted by it anymore, but was also impressed with the architecture design of the process.  </p>\n<p>This is likely why peoples code is failing a lot -- because during inference, they run out of memory.  I would suggest for folks to run inference on the train set because we were told that the test set is around the same size (90% of the size I think).</p>\n<p>(this is my first competition where I have submitted more than one solution (more than just copy/submit notebook) … this might be normal for other competitions, but it was a surprise for me… thanks!)</p>",
      "rawMarkdown": "When we are satisfied with a notebook, we hit Submit ... we can then watch it run, see the logs, etc (kind of exciting to watch), but that run that we see is just a ruse and is only their help us to know whether the code had any issues.  And to my surprise, that run is still only using the 10 test set row dataset.\n\nThere must be another hidden / phantom run that we are unable to see the details for, correct?  That is what is actually scoring our run.  \n\nThis is a pretty clever way to hide the test set from us.  I was super apprehensive about exporting the test set from my \"submit\" code, thinking it was illegal to do so.  I reluctantly did a \"shape\" on it and saw it was only 10 rows.  Was relieved not to be tempted by it anymore, but was also impressed with the architecture design of the process.  \n\nThis is likely why peoples code is failing a lot -- because during inference, they run out of memory.  I would suggest for folks to run inference on the train set because we were told that the test set is around the same size (90% of the size I think).\n\n(this is my first competition where I have submitted more than one solution (more than just copy/submit notebook) ... this might be normal for other competitions, but it was a surprise for me... thanks!)",
      "votes": null
    },
    {
      "id": "2715748",
      "postDate": "03/25/2024 16:51:56",
      "content": "<p>Yes, this is a code competition where you submit your code and not the final scores. Showing you the log output would result in leaking the dataset in a matter of days. </p>",
      "rawMarkdown": "Yes, this is a code competition where you submit your code and not the final scores. Showing you the log output would result in leaking the dataset in a matter of days.",
      "votes": null
    },
    {
      "id": "2716231",
      "postDate": "03/25/2024 22:35:24",
      "content": "<p>I think that when we hit Submit the code runs using the real full test set, you can check the time the notebook needs to read, process and predict on train set is about the same it takes to score because the real test is about the same size with the train set. So I supose kaggle has already the scores for the private leaderboard, I doubt that they are using just 30% of the full test now and then re-run the notebook on the rest.</p>",
      "rawMarkdown": "I think that when we hit Submit the code runs using the real full test set, you can check the time the notebook needs to read, process and predict on train set is about the same it takes to score because the real test is about the same size with the train set. So I supose kaggle has already the scores for the private leaderboard, I doubt that they are using just 30% of the full test now and then re-run the notebook on the rest.",
      "votes": null
    },
    {
      "id": "2716313",
      "postDate": "03/26/2024 00:50:59",
      "content": "<p>Yes and no.  </p>\n<p>It IS running the code on the 10 test rows and gives us visibility to that run (you can test it yourself by returning the shape of your test data) … and then the hidden (phantom) run is something we are not able to see because it is running on the full test set … sometimes when my \"visible\" job finishes, my score does not post for another 5+ minutes … I think that is because it is still creating the \"real\" submission file.  </p>",
      "rawMarkdown": "Yes and no.  \n\nIt IS running the code on the 10 test rows and gives us visibility to that run (you can test it yourself by returning the shape of your test data) ... and then the hidden (phantom) run is something we are not able to see because it is running on the full test set ... sometimes when my \"visible\" job finishes, my score does not post for another 5+ minutes ... I think that is because it is still creating the \"real\" submission file.",
      "votes": null
    },
    {
      "id": "2716672",
      "postDate": "03/26/2024 06:43:49",
      "content": "<p>Yes, it was my point too - the full test is submited in the background and in my case it takes more than 1 hour because the training set processing in my case takes also so long :)</p>",
      "rawMarkdown": "Yes, it was my point too - the full test is submited in the background and in my case it takes more than 1 hour because the training set processing in my case takes also so long :)",
      "votes": null
    },
    {
      "id": "2768986",
      "postDate": "04/23/2024 05:34:24",
      "content": "<blockquote>\n  <p>Showing you the log output would result in leaking the dataset in a matter of days.<br>\n  When I saw the logs, that was my first thought too 😆. I even see that there is a github container that you can download and run and files like <strong>results</strong>.html</p>\n</blockquote>\n<p>I mean I was curious more than anything to know how it gets swapped out ( do you guys keep the variables same or its how kaggle inserts the comp dataset and you just replace that ) </p>\n<p>As a software dev, it was really interesting to think about that ( when i first saw the logs) </p>",
      "rawMarkdown": "> Showing you the log output would result in leaking the dataset in a matter of days.\nWhen I saw the logs, that was my first thought too 😆. I even see that there is a github container that you can download and run and files like __results__.html\n\nI mean I was curious more than anything to know how it gets swapped out ( do you guys keep the variables same or its how kaggle inserts the comp dataset and you just replace that ) \n\nAs a software dev, it was really interesting to think about that ( when i first saw the logs)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2715748,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "03/25/2024 16:51:56",
      "content": "<p>Yes, this is a code competition where you submit your code and not the final scores. Showing you the log output would result in leaking the dataset in a matter of days. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2768986,
          "author_name": "luciferisback",
          "author_url": "",
          "post_date": "04/23/2024 05:34:24",
          "content": "<blockquote>\n  <p>Showing you the log output would result in leaking the dataset in a matter of days.<br>\n  When I saw the logs, that was my first thought too 😆. I even see that there is a github container that you can download and run and files like <strong>results</strong>.html</p>\n</blockquote>\n<p>I mean I was curious more than anything to know how it gets swapped out ( do you guys keep the variables same or its how kaggle inserts the comp dataset and you just replace that ) </p>\n<p>As a software dev, it was really interesting to think about that ( when i first saw the logs) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2716231,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "03/25/2024 22:35:24",
      "content": "<p>I think that when we hit Submit the code runs using the real full test set, you can check the time the notebook needs to read, process and predict on train set is about the same it takes to score because the real test is about the same size with the train set. So I supose kaggle has already the scores for the private leaderboard, I doubt that they are using just 30% of the full test now and then re-run the notebook on the rest.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2716313,
          "author_name": "romandovega",
          "author_url": "",
          "post_date": "03/26/2024 00:50:59",
          "content": "<p>Yes and no.  </p>\n<p>It IS running the code on the 10 test rows and gives us visibility to that run (you can test it yourself by returning the shape of your test data) … and then the hidden (phantom) run is something we are not able to see because it is running on the full test set … sometimes when my \"visible\" job finishes, my score does not post for another 5+ minutes … I think that is because it is still creating the \"real\" submission file.  </p>",
          "votes": null,
          "replies": [
            {
              "id": 2716672,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "03/26/2024 06:43:49",
              "content": "<p>Yes, it was my point too - the full test is submited in the background and in my case it takes more than 1 hour because the training set processing in my case takes also so long :)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2715428": "When we are satisfied with a notebook, we hit Submit ... we can then watch it run, see the logs, etc (kind of exciting to watch), but that run that we see is just a ruse and is only their help us to know whether the code had any issues.  And to my surprise, that run is still only using the 10 test set row dataset.\n\nThere must be another hidden / phantom run that we are unable to see the details for, correct?  That is what is actually scoring our run.  \n\nThis is a pretty clever way to hide the test set from us.  I was super apprehensive about exporting the test set from my \"submit\" code, thinking it was illegal to do so.  I reluctantly did a \"shape\" on it and saw it was only 10 rows.  Was relieved not to be tempted by it anymore, but was also impressed with the architecture design of the process.  \n\nThis is likely why peoples code is failing a lot -- because during inference, they run out of memory.  I would suggest for folks to run inference on the train set because we were told that the test set is around the same size (90% of the size I think).\n\n(this is my first competition where I have submitted more than one solution (more than just copy/submit notebook) ... this might be normal for other competitions, but it was a surprise for me... thanks!)",
    "2715748": "Yes, this is a code competition where you submit your code and not the final scores. Showing you the log output would result in leaking the dataset in a matter of days.",
    "2716231": "I think that when we hit Submit the code runs using the real full test set, you can check the time the notebook needs to read, process and predict on train set is about the same it takes to score because the real test is about the same size with the train set. So I supose kaggle has already the scores for the private leaderboard, I doubt that they are using just 30% of the full test now and then re-run the notebook on the rest.",
    "2716313": "Yes and no.  \n\nIt IS running the code on the 10 test rows and gives us visibility to that run (you can test it yourself by returning the shape of your test data) ... and then the hidden (phantom) run is something we are not able to see because it is running on the full test set ... sometimes when my \"visible\" job finishes, my score does not post for another 5+ minutes ... I think that is because it is still creating the \"real\" submission file.",
    "2716672": "Yes, it was my point too - the full test is submited in the background and in my case it takes more than 1 hour because the training set processing in my case takes also so long :)",
    "2768986": "> Showing you the log output would result in leaking the dataset in a matter of days.\nWhen I saw the logs, that was my first thought too 😆. I even see that there is a github container that you can download and run and files like __results__.html\n\nI mean I was curious more than anything to know how it gets swapped out ( do you guys keep the variables same or its how kaggle inserts the comp dataset and you just replace that ) \n\nAs a software dev, it was really interesting to think about that ( when i first saw the logs)"
  },
  "source": "meta"
}