{
  "id": 436977,
  "title": "Anticipate a decent shakeup",
  "url": "/competitions/predict-ai-model-runtime/discussion/436977",
  "author_name": "Peiyuan Liao",
  "post_date": "2023-09-04T21:08:07.257000",
  "votes": 18,
  "comment_count": 19,
  "views": 0,
  "content": "<p>There's only a few examples (&lt; 30) per collection for layout. Assuming that each collection has its own model, this implies that the leaderboard is calculated only using 4-8 examples for your model. Very easy to have huge shakeups since the models are (from an absolute perspective) really bad. If you look at the ranks across different training sessions they are wildly different, especially for the even harder random collections. In addition, the baseline models are already pretty good at tiles compared to layout (&lt; 5% top 5 err from the paper), so we won't see huge differentiation in what will make up 50% of your final score. </p>\n<p>Another perspective to consider: depending on your model architecture, your model may be really good at one type of BERT and really bad at another type of BERT. If the public / private split happens to only include some types of architectures and not others, it will result in some shakeup.</p>\n<p>The takeaway? I guess we have to take the public leaderboard with a grain of salt :)</p>",
  "messages": [
    {
      "id": 2423878,
      "postDate": "2023-09-04T21:08:07.257Z",
      "content": "<p>There's only a few examples (&lt; 30) per collection for layout. Assuming that each collection has its own model, this implies that the leaderboard is calculated only using 4-8 examples for your model. Very easy to have huge shakeups since the models are (from an absolute perspective) really bad. If you look at the ranks across different training sessions they are wildly different, especially for the even harder random collections. In addition, the baseline models are already pretty good at tiles compared to layout (&lt; 5% top 5 err from the paper), so we won't see huge differentiation in what will make up 50% of your final score. </p>\n<p>Another perspective to consider: depending on your model architecture, your model may be really good at one type of BERT and really bad at another type of BERT. If the public / private split happens to only include some types of architectures and not others, it will result in some shakeup.</p>\n<p>The takeaway? I guess we have to take the public leaderboard with a grain of salt :)</p>",
      "rawMarkdown": "There's only a few examples (< 30) per collection for layout. Assuming that each collection has its own model, this implies that the leaderboard is calculated only using 4-8 examples for your model. Very easy to have huge shakeups since the models are (from an absolute perspective) really bad. If you look at the ranks across different training sessions they are wildly different, especially for the even harder random collections. In addition, the baseline models are already pretty good at tiles compared to layout (< 5% top 5 err from the paper), so we won't see huge differentiation in what will make up 50% of your final score. \n\nAnother perspective to consider: depending on your model architecture, your model may be really good at one type of BERT and really bad at another type of BERT. If the public / private split happens to only include some types of architectures and not others, it will result in some shakeup.\n\nThe takeaway? I guess we have to take the public leaderboard with a grain of salt :)",
      "votes": 16
    },
    {
      "id": 2424168,
      "postDate": "2023-09-05T04:56:14.250Z",
      "content": "<p>That is an interesting perspective - thanks for sharing.</p>\n<p>I felt the same based on the discrepancy between the validation score and the public LB score.</p>\n<p>My LGBM tile model scores:</p>\n<ul>\n<li>Local validation 0.86</li>\n<li>Public LB estimation: 0.247 </li>\n</ul>\n<p>The absolute discrepancy between validation and LB is not necessarily bad if the two correlate, but - 0.86 vs 0.25 seems too much</p>",
      "rawMarkdown": "That is an interesting perspective - thanks for sharing.\n\nI felt the same based on the discrepancy between the validation score and the public LB score.\n\nMy LGBM tile model scores:\n- Local validation 0.86\n- Public LB estimation: 0.247 \n\nThe absolute discrepancy between validation and LB is not necessarily bad if the two correlate, but - 0.86 vs 0.25 seems too much",
      "votes": 4,
      "replies": [
        {
          "id": 2424449,
          "postDate": "2023-09-05T09:34:06.173Z",
          "content": "<p>Aren't the scores dependent on the number of data points in the Public LB for tile? I mean if tile cases only take up 20% of all cases, it makes sense to have a much lower score if your model only predicts tile cases.</p>",
          "rawMarkdown": "Aren't the scores dependent on the number of data points in the Public LB for tile? I mean if tile cases only take up 20% of all cases, it makes sense to have a much lower score if your model only predicts tile cases.",
          "replies": [
            {
              "id": 2425281,
              "postDate": "2023-09-05T18:52:28.687Z",
              "content": "<p>Yes, each collection contributes to 20% of the final score, so tile collection is just 20%.</p>",
              "rawMarkdown": "Yes, each collection contributes to 20% of the final score, so tile collection is just 20%.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2424735,
          "postDate": "2023-09-05T13:11:54.937Z",
          "content": "<p>I think your tile local score is consistent with tile your public score.<br>\n\"There are 5 data collections in total:layout:xla:random, layout:xla:default, layout:nlp:random, layout:nlp:default, and tile:xla. The final score will be the average across all collections.\"<br>\nTherefore, tile score accounts for 20% of the total score.<br>\n0.86 * 0.2 = 0.172, which is similar to 0.183.</p>",
          "rawMarkdown": "I think your tile local score is consistent with tile your public score.\n\"There are 5 data collections in total:layout:xla:random, layout:xla:default, layout:nlp:random, layout:nlp:default, and tile:xla. The final score will be the average across all collections.\"\nTherefore, tile score accounts for 20% of the total score.\n0.86 * 0.2 = 0.172, which is similar to 0.183.\n",
          "votes": 1,
          "replies": [
            {
              "id": 2424792,
              "postDate": "2023-09-05T13:49:54.153Z",
              "content": "<blockquote>\n  <p>As driven by realistic requirements, we use two evaluation metrics, and average them.</p>\n</blockquote>\n<p>I thought the tile contributes to 50% score, no?</p>",
              "rawMarkdown": ">As driven by realistic requirements, we use two evaluation metrics, and average them.\n\nI thought the tile contributes to 50% score, no?"
            },
            {
              "id": 2424814,
              "postDate": "2023-09-05T14:08:19.307Z",
              "content": "<p>I saw \"As driven by realistic requirements, we use two evaluation metrics, and average them.\", and I thought title contributes to 50% score.<br>\nToday, I saw \"There are 5 data collections in total:layout:xla:random, layout:xla:default, layout:nlp:random, layout:nlp:default, and tile:xla. The final score will be the average across all collections.\", and I guess title contributes to 20% score. In this situation, local performance seems to be consist with online performance.</p>",
              "rawMarkdown": "I saw \"As driven by realistic requirements, we use two evaluation metrics, and average them.\", and I thought title contributes to 50% score.\nToday, I saw \"There are 5 data collections in total:layout:xla:random, layout:xla:default, layout:nlp:random, layout:nlp:default, and tile:xla. The final score will be the average across all collections.\", and I guess title contributes to 20% score. In this situation, local performance seems to be consist with online performance."
            },
            {
              "id": 2424819,
              "postDate": "2023-09-05T14:13:00.357Z",
              "content": "<p>It would seem so - with 0 on layout collections, and 0.86 on 20% related to tile, I should get 0.172 as you pointed out.<br>\nBut this means in the sample submission, layout collections score 0, while XLA scores around 0.55 (just by picking 0;1;2;3;4)<br>\nI need to verify if picking 0;1;2;3;4 scores 0.55 on local validation</p>",
              "rawMarkdown": "It would seem so - with 0 on layout collections, and 0.86 on 20% related to tile, I should get 0.172 as you pointed out.\nBut this means in the sample submission, layout collections score 0, while XLA scores around 0.55 (just by picking 0;1;2;3;4)\nI need to verify if picking 0;1;2;3;4 scores 0.55 on local validation"
            },
            {
              "id": 2425358,
              "postDate": "2023-09-05T20:29:58.747Z",
              "content": "<p>My local stats correlate reasonably well with the leaderboard score. How are you performing validation?</p>",
              "rawMarkdown": "My local stats correlate reasonably well with the leaderboard score. How are you performing validation?"
            },
            {
              "id": 2425405,
              "postDate": "2023-09-05T21:41:19.537Z",
              "content": "<p>We average 5 collections: \"tiles:xla\", \"layout:nlp:default\", \"layout:nlp:random\", \"layout:xla:default\", \"layout:xla:random\". Each contributing 1/5</p>",
              "rawMarkdown": "We average 5 collections: \"tiles:xla\", \"layout:nlp:default\", \"layout:nlp:random\", \"layout:xla:default\", \"layout:xla:random\". Each contributing 1/5",
              "votes": 2
            },
            {
              "id": 2425778,
              "postDate": "2023-09-06T07:11:52.070Z",
              "content": "<p>Yes,I think I’ll give the accurate estimate of what we are expecting </p>",
              "rawMarkdown": "Yes,I think I’ll give the accurate estimate of what we are expecting "
            },
            {
              "id": 2426183,
              "postDate": "2023-09-06T13:09:59.797Z",
              "content": "<p>in \"the tiles：xla\" dataset i get 0.162, in the \" tiles：xla\" i get 0.172, Combined they should get 0.3+. But now my score is only 0.19, is this a normal performance？</p>",
              "rawMarkdown": "in \"the tiles：xla\" dataset i get 0.162, in the \" tiles：xla\" i get 0.172, Combined they should get 0.3+. But now my score is only 0.19, is this a normal performance？"
            },
            {
              "id": 2427639,
              "postDate": "2023-09-07T11:11:26.433Z",
              "content": "<p>Hello senior,<br>\nWhen I ran it locally, I also encountered a similar issue. How did you resolve it? Or is there something I might be misunderstanding? Thank you.</p>\n<p>学长您好，我在本地端运行的时候也出现了类似的问题，请问您是如何解决的那？或者这样的理解有什么地方不正确那？谢谢您</p>",
              "rawMarkdown": "Hello senior,\nWhen I ran it locally, I also encountered a similar issue. How did you resolve it? Or is there something I might be misunderstanding? Thank you.\n\n学长您好，我在本地端运行的时候也出现了类似的问题，请问您是如何解决的那？或者这样的理解有什么地方不正确那？谢谢您",
              "isDeleted": true
            },
            {
              "id": 2476513,
              "postDate": "2023-10-10T17:15:22.130Z",
              "content": "<p>Hi Sami,<br>\nAs you clarified, the score is computed by averaging the 5 collections' scores. if I submit collections individually in separate submissions (assuming the scores for all the other fours are zero), the score provided in PL seems to be correct (=1/5 contribution). However, if I submit them collectively in sisingle submission, the scores won't follow the same average approach. Given collection 1 score and collection 2 score (both provided by PL), the result on PL for both of them submitted in one file won't be the summation of the scores. Am I missing something? Any clarification is greatly appreciated.</p>",
              "rawMarkdown": "Hi Sami,\nAs you clarified, the score is computed by averaging the 5 collections' scores. if I submit collections individually in separate submissions (assuming the scores for all the other fours are zero), the score provided in PL seems to be correct (=1/5 contribution). However, if I submit them collectively in sisingle submission, the scores won't follow the same average approach. Given collection 1 score and collection 2 score (both provided by PL), the result on PL for both of them submitted in one file won't be the summation of the scores. Am I missing something? Any clarification is greatly appreciated."
            },
            {
              "id": 2476619,
              "postDate": "2023-10-10T19:07:15.737Z",
              "content": "<p>Hey MiHu, Did you figure out how the combination of scores works? </p>",
              "rawMarkdown": "Hey MiHu, Did you figure out how the combination of scores works? "
            },
            {
              "id": 2476655,
              "postDate": "2023-10-10T19:52:54.090Z",
              "content": "<p>If your submission contains predictions for only one collection, our score calculation will use deterministically random predictions for the rest, so the scores on the missing collections won't be zero. Therefore, when you submit predictions of all collections together, the score won't be the summation.</p>",
              "rawMarkdown": "If your submission contains predictions for only one collection, our score calculation will use deterministically random predictions for the rest, so the scores on the missing collections won't be zero. Therefore, when you submit predictions of all collections together, the score won't be the summation.",
              "votes": 1
            },
            {
              "id": 2476667,
              "postDate": "2023-10-10T20:13:27.767Z",
              "content": "<p>Thanks so much for the reply! That makes sense.</p>",
              "rawMarkdown": "Thanks so much for the reply! That makes sense."
            }
          ]
        }
      ]
    },
    {
      "id": 2425407,
      "postDate": "2023-09-05T21:43:48.727Z",
      "content": "<p>For what its worth, currently, the top 6 teams on the public leaderboard, also the top 6 teams on the private leaderboard. Therefore, there is some decent correlation between private and public leaderboard. Of course, if two public submissions are within a tiny fraction away from one another, there is a chance they will flip positions on the private leaderboard.</p>",
      "rawMarkdown": "For what its worth, currently, the top 6 teams on the public leaderboard, also the top 6 teams on the private leaderboard. Therefore, there is some decent correlation between private and public leaderboard. Of course, if two public submissions are within a tiny fraction away from one another, there is a chance they will flip positions on the private leaderboard.",
      "votes": -4,
      "replies": [
        {
          "id": 2425568,
          "postDate": "2023-09-06T03:40:57.137Z",
          "content": "<p>I understand your intent might be to provide transparency, but please keep in mind sharing this kind of info about the private leaderboard inadvertently influences participants' strategies. It's really against the whole idea of these types of competitions, where the goal is to create models that generalize well to unseen data. I think of it as similar to sharing with participants in a double blind trial if they recieved a placebo….</p>",
          "rawMarkdown": "I understand your intent might be to provide transparency, but please keep in mind sharing this kind of info about the private leaderboard inadvertently influences participants' strategies. It's really against the whole idea of these types of competitions, where the goal is to create models that generalize well to unseen data. I think of it as similar to sharing with participants in a double blind trial if they recieved a placebo....",
          "votes": 26
        }
      ]
    },
    {
      "id": 2427079,
      "postDate": "2023-09-07T04:44:59.760Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2424168,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2023-09-05T04:56:14.250000",
      "content": "<p>That is an interesting perspective - thanks for sharing.</p>\n<p>I felt the same based on the discrepancy between the validation score and the public LB score.</p>\n<p>My LGBM tile model scores:</p>\n<ul>\n<li>Local validation 0.86</li>\n<li>Public LB estimation: 0.247 </li>\n</ul>\n<p>The absolute discrepancy between validation and LB is not necessarily bad if the two correlate, but - 0.86 vs 0.25 seems too much</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2424449,
          "author_name": "PassengerC07",
          "author_url": "",
          "post_date": "2023-09-05T09:34:06.173000",
          "content": "<p>Aren't the scores dependent on the number of data points in the Public LB for tile? I mean if tile cases only take up 20% of all cases, it makes sense to have a much lower score if your model only predicts tile cases.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2425281,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-09-05T18:52:28.687000",
              "content": "<p>Yes, each collection contributes to 20% of the final score, so tile collection is just 20%.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2424735,
          "author_name": "Bruce",
          "author_url": "",
          "post_date": "2023-09-05T13:11:54.937000",
          "content": "<p>I think your tile local score is consistent with tile your public score.<br>\n\"There are 5 data collections in total:layout:xla:random, layout:xla:default, layout:nlp:random, layout:nlp:default, and tile:xla. The final score will be the average across all collections.\"<br>\nTherefore, tile score accounts for 20% of the total score.<br>\n0.86 * 0.2 = 0.172, which is similar to 0.183.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2424792,
              "author_name": "narsil (jobs-in-data.com)",
              "author_url": "",
              "post_date": "2023-09-05T13:49:54.153000",
              "content": "<blockquote>\n  <p>As driven by realistic requirements, we use two evaluation metrics, and average them.</p>\n</blockquote>\n<p>I thought the tile contributes to 50% score, no?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2424814,
              "author_name": "Bruce",
              "author_url": "",
              "post_date": "2023-09-05T14:08:19.307000",
              "content": "<p>I saw \"As driven by realistic requirements, we use two evaluation metrics, and average them.\", and I thought title contributes to 50% score.<br>\nToday, I saw \"There are 5 data collections in total:layout:xla:random, layout:xla:default, layout:nlp:random, layout:nlp:default, and tile:xla. The final score will be the average across all collections.\", and I guess title contributes to 20% score. In this situation, local performance seems to be consist with online performance.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2424819,
              "author_name": "narsil (jobs-in-data.com)",
              "author_url": "",
              "post_date": "2023-09-05T14:13:00.357000",
              "content": "<p>It would seem so - with 0 on layout collections, and 0.86 on 20% related to tile, I should get 0.172 as you pointed out.<br>\nBut this means in the sample submission, layout collections score 0, while XLA scores around 0.55 (just by picking 0;1;2;3;4)<br>\nI need to verify if picking 0;1;2;3;4 scores 0.55 on local validation</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2425358,
              "author_name": "Peiyuan Liao",
              "author_url": "",
              "post_date": "2023-09-05T20:29:58.747000",
              "content": "<p>My local stats correlate reasonably well with the leaderboard score. How are you performing validation?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2425405,
              "author_name": "Sami Abu-El-Haija",
              "author_url": "",
              "post_date": "2023-09-05T21:41:19.537000",
              "content": "<p>We average 5 collections: \"tiles:xla\", \"layout:nlp:default\", \"layout:nlp:random\", \"layout:xla:default\", \"layout:xla:random\". Each contributing 1/5</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2425778,
              "author_name": "Abdulkadir Aliyu",
              "author_url": "",
              "post_date": "2023-09-06T07:11:52.070000",
              "content": "<p>Yes,I think I’ll give the accurate estimate of what we are expecting </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2426183,
              "author_name": "MiHu",
              "author_url": "",
              "post_date": "2023-09-06T13:09:59.797000",
              "content": "<p>in \"the tiles：xla\" dataset i get 0.162, in the \" tiles：xla\" i get 0.172, Combined they should get 0.3+. But now my score is only 0.19, is this a normal performance？</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2427639,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-07T11:11:26.433000",
              "content": "<p>Hello senior,<br>\nWhen I ran it locally, I also encountered a similar issue. How did you resolve it? Or is there something I might be misunderstanding? Thank you.</p>\n<p>学长您好，我在本地端运行的时候也出现了类似的问题，请问您是如何解决的那？或者这样的理解有什么地方不正确那？谢谢您</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2476513,
              "author_name": "Ali Jazayeri",
              "author_url": "",
              "post_date": "2023-10-10T17:15:22.130000",
              "content": "<p>Hi Sami,<br>\nAs you clarified, the score is computed by averaging the 5 collections' scores. if I submit collections individually in separate submissions (assuming the scores for all the other fours are zero), the score provided in PL seems to be correct (=1/5 contribution). However, if I submit them collectively in sisingle submission, the scores won't follow the same average approach. Given collection 1 score and collection 2 score (both provided by PL), the result on PL for both of them submitted in one file won't be the summation of the scores. Am I missing something? Any clarification is greatly appreciated.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2476619,
              "author_name": "Ali Jazayeri",
              "author_url": "",
              "post_date": "2023-10-10T19:07:15.737000",
              "content": "<p>Hey MiHu, Did you figure out how the combination of scores works? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2476655,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-10-10T19:52:54.090000",
              "content": "<p>If your submission contains predictions for only one collection, our score calculation will use deterministically random predictions for the rest, so the scores on the missing collections won't be zero. Therefore, when you submit predictions of all collections together, the score won't be the summation.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2476667,
              "author_name": "Ali Jazayeri",
              "author_url": "",
              "post_date": "2023-10-10T20:13:27.767000",
              "content": "<p>Thanks so much for the reply! That makes sense.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2425407,
      "author_name": "Sami Abu-El-Haija",
      "author_url": "",
      "post_date": "2023-09-05T21:43:48.727000",
      "content": "<p>For what its worth, currently, the top 6 teams on the public leaderboard, also the top 6 teams on the private leaderboard. Therefore, there is some decent correlation between private and public leaderboard. Of course, if two public submissions are within a tiny fraction away from one another, there is a chance they will flip positions on the private leaderboard.</p>",
      "votes": -4,
      "replies": [
        {
          "id": 2425568,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2023-09-06T03:40:57.137000",
          "content": "<p>I understand your intent might be to provide transparency, but please keep in mind sharing this kind of info about the private leaderboard inadvertently influences participants' strategies. It's really against the whole idea of these types of competitions, where the goal is to create models that generalize well to unseen data. I think of it as similar to sharing with participants in a double blind trial if they recieved a placebo….</p>",
          "votes": 26,
          "replies": []
        }
      ]
    },
    {
      "id": 2427079,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-07T04:44:59.760000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2423878": "There's only a few examples (< 30) per collection for layout. Assuming that each collection has its own model, this implies that the leaderboard is calculated only using 4-8 examples for your model. Very easy to have huge shakeups since the models are (from an absolute perspective) really bad. If you look at the ranks across different training sessions they are wildly different, especially for the even harder random collections. In addition, the baseline models are already pretty good at tiles compared to layout (< 5% top 5 err from the paper), so we won't see huge differentiation in what will make up 50% of your final score. \n\nAnother perspective to consider: depending on your model architecture, your model may be really good at one type of BERT and really bad at another type of BERT. If the public / private split happens to only include some types of architectures and not others, it will result in some shakeup.\n\nThe takeaway? I guess we have to take the public leaderboard with a grain of salt :)",
    "2424168": "That is an interesting perspective - thanks for sharing.\n\nI felt the same based on the discrepancy between the validation score and the public LB score.\n\nMy LGBM tile model scores:\n- Local validation 0.86\n- Public LB estimation: 0.247 \n\nThe absolute discrepancy between validation and LB is not necessarily bad if the two correlate, but - 0.86 vs 0.25 seems too much",
    "2425407": "For what its worth, currently, the top 6 teams on the public leaderboard, also the top 6 teams on the private leaderboard. Therefore, there is some decent correlation between private and public leaderboard. Of course, if two public submissions are within a tiny fraction away from one another, there is a chance they will flip positions on the private leaderboard.",
    "2427079": ""
  }
}