{
  "id": 505936,
  "title": "How is the public score calculated?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/505936",
  "author_name": "",
  "post_date": "2024-05-19T19:14:45.036035600Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The test set consists of two parts: public test and private test. I can see several possible options how a submission is scored:</p>\n<p>1) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits one regression to compute the coefficients a and b (of the linear regression, a * x + b, which is fit through the weekly gini scores). After that, the public score is calculated using the values of a and b, and the private score is calculated using the same values of a and b. It would be strange, because it means some sort of leakage from the private test set.</p>\n<p>2) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits two regressions to compute two pairs of coefficients: a_public, b_public, a_private, b_private. Then, a_public and b_public are used to calculate the public score. Then, a_private and b_private are used to calculate the private score.</p>\n<p>3) A submitted notebook receives only public test set, fits a regression to compute a_public and b_public, and the public score is computed. After that, the submitted notebook gets the private test set, fits another regression to compute a_private and b_private, and the private score is finally computed.</p>\n<p>4) Something else?</p>\n<p>Which one is correct?</p>",
  "messages": [
    {
      "id": "2824453",
      "postDate": "05/19/2024 19:14:45",
      "content": "<p>The test set consists of two parts: public test and private test. I can see several possible options how a submission is scored:</p>\n<p>1) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits one regression to compute the coefficients a and b (of the linear regression, a * x + b, which is fit through the weekly gini scores). After that, the public score is calculated using the values of a and b, and the private score is calculated using the same values of a and b. It would be strange, because it means some sort of leakage from the private test set.</p>\n<p>2) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits two regressions to compute two pairs of coefficients: a_public, b_public, a_private, b_private. Then, a_public and b_public are used to calculate the public score. Then, a_private and b_private are used to calculate the private score.</p>\n<p>3) A submitted notebook receives only public test set, fits a regression to compute a_public and b_public, and the public score is computed. After that, the submitted notebook gets the private test set, fits another regression to compute a_private and b_private, and the private score is finally computed.</p>\n<p>4) Something else?</p>\n<p>Which one is correct?</p>",
      "rawMarkdown": "The test set consists of two parts: public test and private test. I can see several possible options how a submission is scored:\n\n1) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits one regression to compute the coefficients a and b (of the linear regression, a * x + b, which is fit through the weekly gini scores). After that, the public score is calculated using the values of a and b, and the private score is calculated using the same values of a and b. It would be strange, because it means some sort of leakage from the private test set.\n\n2) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits two regressions to compute two pairs of coefficients: a_public, b_public, a_private, b_private. Then, a_public and b_public are used to calculate the public score. Then, a_private and b_private are used to calculate the private score.\n\n3) A submitted notebook receives only public test set, fits a regression to compute a_public and b_public, and the public score is computed. After that, the submitted notebook gets the private test set, fits another regression to compute a_private and b_private, and the private score is finally computed.\n\n4) Something else?\n\nWhich one is correct?",
      "votes": null
    },
    {
      "id": "2824620",
      "postDate": "05/19/2024 23:28:30",
      "content": "<p>I also want to know this question.</p>",
      "rawMarkdown": "I also want to know this question.",
      "votes": null
    },
    {
      "id": "2824777",
      "postDate": "05/20/2024 03:19:44",
      "content": "<p>I am guessing it is (2) or (3) so that LB probing has no effect.  I believe it is (2)</p>",
      "rawMarkdown": "I am guessing it is (2) or (3) so that LB probing has no effect.  I believe it is (2)",
      "votes": null
    },
    {
      "id": "2827187",
      "postDate": "05/21/2024 09:34:35",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> any insight?</p>",
      "rawMarkdown": "tomasjeline2 any insight?",
      "votes": null
    },
    {
      "id": "2833721",
      "postDate": "05/24/2024 11:16:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> ,<br>\nthis is more question to Kaggle, I am not 100% sure but I think it's option 2 based on discussion we had with Kaggle.<br>\nNotebook receives full dataset and produces score for each record in dataset. Only after that is done split to public/private, and run evaluation script (fit regression and computes a,b coefficients).<br>\nBut as I said, I don't know details about Kaggle's \"backend\", so I might be wrong…</p>",
      "rawMarkdown": "Hi @julianmukaj ,\nthis is more question to Kaggle, I am not 100% sure but I think it's option 2 based on discussion we had with Kaggle.\nNotebook receives full dataset and produces score for each record in dataset. Only after that is done split to public/private, and run evaluation script (fit regression and computes a,b coefficients).\nBut as I said, I don't know details about Kaggle's \"backend\", so I might be wrong...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2824620,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "05/19/2024 23:28:30",
      "content": "<p>I also want to know this question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2824777,
      "author_name": "fadylabib",
      "author_url": "",
      "post_date": "05/20/2024 03:19:44",
      "content": "<p>I am guessing it is (2) or (3) so that LB probing has no effect.  I believe it is (2)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2827187,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "05/21/2024 09:34:35",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> any insight?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2833721,
          "author_name": "tomasjeline2",
          "author_url": "",
          "post_date": "05/24/2024 11:16:19",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> ,<br>\nthis is more question to Kaggle, I am not 100% sure but I think it's option 2 based on discussion we had with Kaggle.<br>\nNotebook receives full dataset and produces score for each record in dataset. Only after that is done split to public/private, and run evaluation script (fit regression and computes a,b coefficients).<br>\nBut as I said, I don't know details about Kaggle's \"backend\", so I might be wrong…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2824453": "The test set consists of two parts: public test and private test. I can see several possible options how a submission is scored:\n\n1) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits one regression to compute the coefficients a and b (of the linear regression, a * x + b, which is fit through the weekly gini scores). After that, the public score is calculated using the values of a and b, and the private score is calculated using the same values of a and b. It would be strange, because it means some sort of leakage from the private test set.\n\n2) A submitted notebook receives the whole test set, predicts for every row of the test set, then fits two regressions to compute two pairs of coefficients: a_public, b_public, a_private, b_private. Then, a_public and b_public are used to calculate the public score. Then, a_private and b_private are used to calculate the private score.\n\n3) A submitted notebook receives only public test set, fits a regression to compute a_public and b_public, and the public score is computed. After that, the submitted notebook gets the private test set, fits another regression to compute a_private and b_private, and the private score is finally computed.\n\n4) Something else?\n\nWhich one is correct?",
    "2824620": "I also want to know this question.",
    "2824777": "I am guessing it is (2) or (3) so that LB probing has no effect.  I believe it is (2)",
    "2827187": "tomasjeline2 any insight?",
    "2833721": "Hi @julianmukaj ,\nthis is more question to Kaggle, I am not 100% sure but I think it's option 2 based on discussion we had with Kaggle.\nNotebook receives full dataset and produces score for each record in dataset. Only after that is done split to public/private, and run evaluation script (fit regression and computes a,b coefficients).\nBut as I said, I don't know details about Kaggle's \"backend\", so I might be wrong..."
  },
  "source": "meta"
}