{
  "id": 261755,
  "title": "MCMAE calculation for PB",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/261755",
  "author_name": "",
  "post_date": "2021-08-05T02:07:20.345493400Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I got involved late in this competition and spent most of my time building the features and fixing the submission errors. By then the updated timeline was out and I used it for my training.</p>\n<p>I was under the impression that the public LB was based on the May data (May 1st to may 31st). Out of curiosity I tested my models against this period, in addition to my validation periods. When I submitted the first model (a few days after I figured out how to fix the submission errors) I was surprised that the PB score was was a lot worse than the one I calculated (by more than 0.1). The function that I built to calculate the score is extremely simple (clip prediction to 0..100 and MAE), but just in case, I exported the predicted and the true values to Excel and verified the calculation was correct.</p>\n<p>I checked various combinations of dates in Excel and none came close to the public LB. Some folks got a 0 in the LB, so the LB calculation is obviously correct. I suppose that either I used the wrong range or misinterpreted what MCMAE means. The text states that you calculate MAE per column and then average the results. if my intuition and algebra are right, that's the same as calculating the MAE overall. In both cases you sum the individual MAEs and divide them by the total. Considering that all target columns have the same length, MCMAE and MAE should be the same.</p>\n<p>Did anybody face the same \"issue\"? For those we got a 0, what range did you use? What am I missing?</p>",
  "messages": [
    {
      "id": "1449916",
      "postDate": "08/05/2021 02:07:20",
      "content": "<p>I got involved late in this competition and spent most of my time building the features and fixing the submission errors. By then the updated timeline was out and I used it for my training.</p>\n<p>I was under the impression that the public LB was based on the May data (May 1st to may 31st). Out of curiosity I tested my models against this period, in addition to my validation periods. When I submitted the first model (a few days after I figured out how to fix the submission errors) I was surprised that the PB score was was a lot worse than the one I calculated (by more than 0.1). The function that I built to calculate the score is extremely simple (clip prediction to 0..100 and MAE), but just in case, I exported the predicted and the true values to Excel and verified the calculation was correct.</p>\n<p>I checked various combinations of dates in Excel and none came close to the public LB. Some folks got a 0 in the LB, so the LB calculation is obviously correct. I suppose that either I used the wrong range or misinterpreted what MCMAE means. The text states that you calculate MAE per column and then average the results. if my intuition and algebra are right, that's the same as calculating the MAE overall. In both cases you sum the individual MAEs and divide them by the total. Considering that all target columns have the same length, MCMAE and MAE should be the same.</p>\n<p>Did anybody face the same \"issue\"? For those we got a 0, what range did you use? What am I missing?</p>",
      "rawMarkdown": "I got involved late in this competition and spent most of my time building the features and fixing the submission errors. By then the updated timeline was out and I used it for my training.\n\nI was under the impression that the public LB was based on the May data (May 1st to may 31st). Out of curiosity I tested my models against this period, in addition to my validation periods. When I submitted the first model (a few days after I figured out how to fix the submission errors) I was surprised that the PB score was was a lot worse than the one I calculated (by more than 0.1). The function that I built to calculate the score is extremely simple (clip prediction to 0..100 and MAE), but just in case, I exported the predicted and the true values to Excel and verified the calculation was correct.\n\nI checked various combinations of dates in Excel and none came close to the public LB. Some folks got a 0 in the LB, so the LB calculation is obviously correct. I suppose that either I used the wrong range or misinterpreted what MCMAE means. The text states that you calculate MAE per column and then average the results. if my intuition and algebra are right, that's the same as calculating the MAE overall. In both cases you sum the individual MAEs and divide them by the total. Considering that all target columns have the same length, MCMAE and MAE should be the same.\n\nDid anybody face the same \"issue\"? For those we got a 0, what range did you use? What am I missing?",
      "votes": null
    },
    {
      "id": "1449950",
      "postDate": "08/05/2021 02:26:44",
      "content": "<p>Did you filter the players to those included for testing? There are 1187 predictions per day.</p>",
      "rawMarkdown": "Did you filter the players to those included for testing? There are 1187 predictions per day.",
      "votes": null
    },
    {
      "id": "1449988",
      "postDate": "08/05/2021 02:42:58",
      "content": "<p>I published a notebook for imitating the grading on LB that may help you: <a href=\"https://www.kaggle.com/zacchaeus/mlb-api-emulator-with-scoring\" target=\"_blank\">https://www.kaggle.com/zacchaeus/mlb-api-emulator-with-scoring</a></p>",
      "rawMarkdown": "I published a notebook for imitating the grading on LB that may help you: https://www.kaggle.com/zacchaeus/mlb-api-emulator-with-scoring",
      "votes": null
    },
    {
      "id": "1451727",
      "postDate": "08/05/2021 12:17:33",
      "content": "<p>Yes, I did.</p>",
      "rawMarkdown": "Yes, I did.",
      "votes": null
    },
    {
      "id": "1451938",
      "postDate": "08/05/2021 13:31:39",
      "content": "<p>Thanks for the response and for sharing the notebook (upvoted). The calculation of the score seems to be the same I'm using, yet I don't get the same results as the LB. Have you compared the values in the LB against the one your notebook calculates?</p>\n<p>I picked one of my models and only found 4 date ranges for which my calculation matched the public LB (considering all ranges of 1 to 61 consecutive days after Apr 1st). However, none of them was even remotely close for the other 7 models I submitted.</p>\n<p>I wonder what I did wrong. The fact that I can't match the value doesn't really matter per se (I didn't base the evaluation of my models on it). The only reason it interests me is that it may be a symptom of something wrong with the scoring I did use for evaluating my models.</p>",
      "rawMarkdown": "Thanks for the response and for sharing the notebook (upvoted). The calculation of the score seems to be the same I'm using, yet I don't get the same results as the LB. Have you compared the values in the LB against the one your notebook calculates?\n\nI picked one of my models and only found 4 date ranges for which my calculation matched the public LB (considering all ranges of 1 to 61 consecutive days after Apr 1st). However, none of them was even remotely close for the other 7 models I submitted.\n\nI wonder what I did wrong. The fact that I can't match the value doesn't really matter per se (I didn't base the evaluation of my models on it). The only reason it interests me is that it may be a symptom of something wrong with the scoring I did use for evaluating my models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1449950,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "08/05/2021 02:26:44",
      "content": "<p>Did you filter the players to those included for testing? There are 1187 predictions per day.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1451727,
          "author_name": "vialactea",
          "author_url": "",
          "post_date": "08/05/2021 12:17:33",
          "content": "<p>Yes, I did.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1449988,
      "author_name": "zacchaeus",
      "author_url": "",
      "post_date": "08/05/2021 02:42:58",
      "content": "<p>I published a notebook for imitating the grading on LB that may help you: <a href=\"https://www.kaggle.com/zacchaeus/mlb-api-emulator-with-scoring\" target=\"_blank\">https://www.kaggle.com/zacchaeus/mlb-api-emulator-with-scoring</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1451938,
          "author_name": "vialactea",
          "author_url": "",
          "post_date": "08/05/2021 13:31:39",
          "content": "<p>Thanks for the response and for sharing the notebook (upvoted). The calculation of the score seems to be the same I'm using, yet I don't get the same results as the LB. Have you compared the values in the LB against the one your notebook calculates?</p>\n<p>I picked one of my models and only found 4 date ranges for which my calculation matched the public LB (considering all ranges of 1 to 61 consecutive days after Apr 1st). However, none of them was even remotely close for the other 7 models I submitted.</p>\n<p>I wonder what I did wrong. The fact that I can't match the value doesn't really matter per se (I didn't base the evaluation of my models on it). The only reason it interests me is that it may be a symptom of something wrong with the scoring I did use for evaluating my models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1449916": "I got involved late in this competition and spent most of my time building the features and fixing the submission errors. By then the updated timeline was out and I used it for my training.\n\nI was under the impression that the public LB was based on the May data (May 1st to may 31st). Out of curiosity I tested my models against this period, in addition to my validation periods. When I submitted the first model (a few days after I figured out how to fix the submission errors) I was surprised that the PB score was was a lot worse than the one I calculated (by more than 0.1). The function that I built to calculate the score is extremely simple (clip prediction to 0..100 and MAE), but just in case, I exported the predicted and the true values to Excel and verified the calculation was correct.\n\nI checked various combinations of dates in Excel and none came close to the public LB. Some folks got a 0 in the LB, so the LB calculation is obviously correct. I suppose that either I used the wrong range or misinterpreted what MCMAE means. The text states that you calculate MAE per column and then average the results. if my intuition and algebra are right, that's the same as calculating the MAE overall. In both cases you sum the individual MAEs and divide them by the total. Considering that all target columns have the same length, MCMAE and MAE should be the same.\n\nDid anybody face the same \"issue\"? For those we got a 0, what range did you use? What am I missing?",
    "1449950": "Did you filter the players to those included for testing? There are 1187 predictions per day.",
    "1449988": "I published a notebook for imitating the grading on LB that may help you: https://www.kaggle.com/zacchaeus/mlb-api-emulator-with-scoring",
    "1451727": "Yes, I did.",
    "1451938": "Thanks for the response and for sharing the notebook (upvoted). The calculation of the score seems to be the same I'm using, yet I don't get the same results as the LB. Have you compared the values in the LB against the one your notebook calculates?\n\nI picked one of my models and only found 4 date ranges for which my calculation matched the public LB (considering all ranges of 1 to 61 consecutive days after Apr 1st). However, none of them was even remotely close for the other 7 models I submitted.\n\nI wonder what I did wrong. The fact that I can't match the value doesn't really matter per se (I didn't base the evaluation of my models on it). The only reason it interests me is that it may be a symptom of something wrong with the scoring I did use for evaluating my models."
  },
  "source": "meta"
}