{
  "id": 548929,
  "title": "Is the submission score computed by date/time and then averaged, or on the whole set?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548929",
  "author_name": "",
  "post_date": "2024-11-29T14:44:01.894007400Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>My first guess was that R-squared is calculated over the whole test set, since it's not otherwise stated and it makes more sense to rely on the zero-mean quality of the target if the sample is large.</p>\n<p>However, since the predictions have to be made by date and time and the R-squared on my local validation is more reflective of the performance on the hidden set if I compute it grouped by date and time, I was wondering whether the submission score may be calculated as average of groups instead of 'globally'.</p>\n<p>Does anyone have any insight on this?</p>",
  "messages": [
    {
      "id": "3058524",
      "postDate": "11/29/2024 14:44:01",
      "content": "<p>My first guess was that R-squared is calculated over the whole test set, since it's not otherwise stated and it makes more sense to rely on the zero-mean quality of the target if the sample is large.</p>\n<p>However, since the predictions have to be made by date and time and the R-squared on my local validation is more reflective of the performance on the hidden set if I compute it grouped by date and time, I was wondering whether the submission score may be calculated as average of groups instead of 'globally'.</p>\n<p>Does anyone have any insight on this?</p>",
      "rawMarkdown": "My first guess was that R-squared is calculated over the whole test set, since it's not otherwise stated and it makes more sense to rely on the zero-mean quality of the target if the sample is large.\n\nHowever, since the predictions have to be made by date and time and the R-squared on my local validation is more reflective of the performance on the hidden set if I compute it grouped by date and time, I was wondering whether the submission score may be calculated as average of groups instead of 'globally'.\n\nDoes anyone have any insight on this?",
      "votes": null
    },
    {
      "id": "3058539",
      "postDate": "11/29/2024 15:08:39",
      "content": "<p>I suugest you to read the overview of the competition. In that the evalutation para has been written that how will the output is evaluated</p>",
      "rawMarkdown": "I suugest you to read the overview of the competition. In that the evalutation para has been written that how will the output is evaluated",
      "votes": null
    },
    {
      "id": "3058616",
      "postDate": "11/29/2024 17:31:18",
      "content": "<p>I did, and although it doesn't mention grouping, I think it may still be possible that there is grouping going on and thus want to make sure.</p>",
      "rawMarkdown": "I did, and although it doesn't mention grouping, I think it may still be possible that there is grouping going on and thus want to make sure.",
      "votes": null
    },
    {
      "id": "3058703",
      "postDate": "11/29/2024 20:12:27",
      "content": "<p>Yes you re-evaluate your code and see what happens next. All the best </p>",
      "rawMarkdown": "Yes you re-evaluate your code and see what happens next. All the best",
      "votes": null
    },
    {
      "id": "3059726",
      "postDate": "12/01/2024 03:10:48",
      "content": "<p>Have you found any official statement about this? I have the same question.</p>",
      "rawMarkdown": "Have you found any official statement about this? I have the same question.",
      "votes": null
    },
    {
      "id": "3059893",
      "postDate": "12/01/2024 06:44:43",
      "content": "<p>On the whole set. <br>\nI submit twice. Change to different time windows, the lb score is very stable and close to ~14, just like performance on the training set.</p>\n<pre><code>counter += 1\npreds = np.zeros(len())\n\n counter == 23456: \n    preds[0] = 1e4 / [][0] ** 0.5\npredictions = test.select(\n        ,\n        pl.lit(preds).(),\n    )\n</code></pre>",
      "rawMarkdown": "On the whole set. \nI submit twice. Change to different time windows, the lb score is very stable and close to ~14, just like performance on the training set.\n```\ncounter += 1\npreds = np.zeros(len(test))\n# if counter == 1: \nif counter == 23456: \n    preds[0] = 1e4 / test['weight'][0] ** 0.5\npredictions = test.select(\n        'row_id',\n        pl.lit(preds).alias('responder_6'),\n    )\n```",
      "votes": null
    },
    {
      "id": "3060872",
      "postDate": "12/02/2024 07:42:00",
      "content": "<p>Thanks for your test! :)</p>",
      "rawMarkdown": "Thanks for your test! :)",
      "votes": null
    },
    {
      "id": "3060874",
      "postDate": "12/02/2024 07:43:38",
      "content": "<p>No, but see <a href=\"https://www.kaggle.com/hydantess\" target=\"_blank\">@hydantess</a>'s test.</p>",
      "rawMarkdown": "No, but see @hydantess's test.",
      "votes": null
    },
    {
      "id": "3061187",
      "postDate": "12/02/2024 12:35:04",
      "content": "<p>Got it, thanks.</p>",
      "rawMarkdown": "Got it, thanks.",
      "votes": null
    },
    {
      "id": "3061555",
      "postDate": "12/02/2024 19:03:17",
      "content": "<p><a href=\"https://www.kaggle.com/hyd\" target=\"_blank\">@hyd</a> what is this testing exactly?</p>",
      "rawMarkdown": "hyd what is this testing exactly?",
      "votes": null
    },
    {
      "id": "3061767",
      "postDate": "12/03/2024 01:12:45",
      "content": "<p>I keep the numerator almost constant. The lb score will fluctuate a lot if computed by date_id&amp;time_id and then averaged.</p>",
      "rawMarkdown": "I keep the numerator almost constant. The lb score will fluctuate a lot if computed by date_id&time_id and then averaged.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3058539,
      "author_name": "sumit08",
      "author_url": "",
      "post_date": "11/29/2024 15:08:39",
      "content": "<p>I suugest you to read the overview of the competition. In that the evalutation para has been written that how will the output is evaluated</p>",
      "votes": null,
      "replies": [
        {
          "id": 3058616,
          "author_name": "jonassend",
          "author_url": "",
          "post_date": "11/29/2024 17:31:18",
          "content": "<p>I did, and although it doesn't mention grouping, I think it may still be possible that there is grouping going on and thus want to make sure.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3058703,
              "author_name": "sumit08",
              "author_url": "",
              "post_date": "11/29/2024 20:12:27",
              "content": "<p>Yes you re-evaluate your code and see what happens next. All the best </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3059726,
      "author_name": "zui0711",
      "author_url": "",
      "post_date": "12/01/2024 03:10:48",
      "content": "<p>Have you found any official statement about this? I have the same question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3060874,
          "author_name": "jonassend",
          "author_url": "",
          "post_date": "12/02/2024 07:43:38",
          "content": "<p>No, but see <a href=\"https://www.kaggle.com/hydantess\" target=\"_blank\">@hydantess</a>'s test.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3061187,
              "author_name": "zui0711",
              "author_url": "",
              "post_date": "12/02/2024 12:35:04",
              "content": "<p>Got it, thanks.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3059893,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "12/01/2024 06:44:43",
      "content": "<p>On the whole set. <br>\nI submit twice. Change to different time windows, the lb score is very stable and close to ~14, just like performance on the training set.</p>\n<pre><code>counter += 1\npreds = np.zeros(len())\n\n counter == 23456: \n    preds[0] = 1e4 / [][0] ** 0.5\npredictions = test.select(\n        ,\n        pl.lit(preds).(),\n    )\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3060872,
          "author_name": "jonassend",
          "author_url": "",
          "post_date": "12/02/2024 07:42:00",
          "content": "<p>Thanks for your test! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3061555,
          "author_name": "jagofc",
          "author_url": "",
          "post_date": "12/02/2024 19:03:17",
          "content": "<p><a href=\"https://www.kaggle.com/hyd\" target=\"_blank\">@hyd</a> what is this testing exactly?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3061767,
              "author_name": "hydantess",
              "author_url": "",
              "post_date": "12/03/2024 01:12:45",
              "content": "<p>I keep the numerator almost constant. The lb score will fluctuate a lot if computed by date_id&amp;time_id and then averaged.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3058524": "My first guess was that R-squared is calculated over the whole test set, since it's not otherwise stated and it makes more sense to rely on the zero-mean quality of the target if the sample is large.\n\nHowever, since the predictions have to be made by date and time and the R-squared on my local validation is more reflective of the performance on the hidden set if I compute it grouped by date and time, I was wondering whether the submission score may be calculated as average of groups instead of 'globally'.\n\nDoes anyone have any insight on this?",
    "3058539": "I suugest you to read the overview of the competition. In that the evalutation para has been written that how will the output is evaluated",
    "3058616": "I did, and although it doesn't mention grouping, I think it may still be possible that there is grouping going on and thus want to make sure.",
    "3058703": "Yes you re-evaluate your code and see what happens next. All the best",
    "3059726": "Have you found any official statement about this? I have the same question.",
    "3059893": "On the whole set. \nI submit twice. Change to different time windows, the lb score is very stable and close to ~14, just like performance on the training set.\n```\ncounter += 1\npreds = np.zeros(len(test))\n# if counter == 1: \nif counter == 23456: \n    preds[0] = 1e4 / test['weight'][0] ** 0.5\npredictions = test.select(\n        'row_id',\n        pl.lit(preds).alias('responder_6'),\n    )\n```",
    "3060872": "Thanks for your test! :)",
    "3060874": "No, but see @hydantess's test.",
    "3061187": "Got it, thanks.",
    "3061555": "hyd what is this testing exactly?",
    "3061767": "I keep the numerator almost constant. The lb score will fluctuate a lot if computed by date_id&time_id and then averaged."
  },
  "source": "meta"
}