{
  "id": 453134,
  "title": "Test data includes private and public LB data?",
  "url": "/competitions/predict-ai-model-runtime/discussion/453134",
  "author_name": "",
  "post_date": "2023-11-05T04:54:34.107571600Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am currently doing Inference in 5 different Notebooks and Combining the Csv in one in end in another Notebook. Is this methodology correct? or is there are some private LB data that I am missing out on?</p>\n<p>Thank You.</p>",
  "messages": [
    {
      "id": "2512981",
      "postDate": "11/05/2023 04:54:34",
      "content": "<p>I am currently doing Inference in 5 different Notebooks and Combining the Csv in one in end in another Notebook. Is this methodology correct? or is there are some private LB data that I am missing out on?</p>\n<p>Thank You.</p>",
      "rawMarkdown": "I am currently doing Inference in 5 different Notebooks and Combining the Csv in one in end in another Notebook. Is this methodology correct? or is there are some private LB data that I am missing out on?\n\nThank You.",
      "votes": null
    },
    {
      "id": "2513245",
      "postDate": "11/05/2023 10:10:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/goelyash\" target=\"_blank\">@goelyash</a>,</p>\n<p>Exactly, the current test data on hand contains both public and private LB data. One thing to note is that the number of graphs used to derive the public score is quite small, roughly 4 for <em>layout-xla-default / random</em> and 8 for <em>layout-nlp-default / random</em>. Considering the small dataset, the public score can be sensitive to randomness even if you just change the random seed and leave all the other params unchanged. For more detailed discussion, you can refer to <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436977\" target=\"_blank\">this great forum</a>.</p>",
      "rawMarkdown": "Hi @goelyash,\n\nExactly, the current test data on hand contains both public and private LB data. One thing to note is that the number of graphs used to derive the public score is quite small, roughly 4 for *layout-xla-default / random* and 8 for *layout-nlp-default / random*. Considering the small dataset, the public score can be sensitive to randomness even if you just change the random seed and leave all the other params unchanged. For more detailed discussion, you can refer to [this great forum](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436977).",
      "votes": null
    },
    {
      "id": "2517341",
      "postDate": "11/08/2023 11:49:31",
      "content": "<p>Correct: the provided test data contains both public and private leaderboard data!</p>",
      "rawMarkdown": "Correct: the provided test data contains both public and private leaderboard data!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2513245,
      "author_name": "abaojiang",
      "author_url": "",
      "post_date": "11/05/2023 10:10:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/goelyash\" target=\"_blank\">@goelyash</a>,</p>\n<p>Exactly, the current test data on hand contains both public and private LB data. One thing to note is that the number of graphs used to derive the public score is quite small, roughly 4 for <em>layout-xla-default / random</em> and 8 for <em>layout-nlp-default / random</em>. Considering the small dataset, the public score can be sensitive to randomness even if you just change the random seed and leave all the other params unchanged. For more detailed discussion, you can refer to <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436977\" target=\"_blank\">this great forum</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2517341,
          "author_name": "samihaija",
          "author_url": "",
          "post_date": "11/08/2023 11:49:31",
          "content": "<p>Correct: the provided test data contains both public and private leaderboard data!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2512981": "I am currently doing Inference in 5 different Notebooks and Combining the Csv in one in end in another Notebook. Is this methodology correct? or is there are some private LB data that I am missing out on?\n\nThank You.",
    "2513245": "Hi @goelyash,\n\nExactly, the current test data on hand contains both public and private LB data. One thing to note is that the number of graphs used to derive the public score is quite small, roughly 4 for *layout-xla-default / random* and 8 for *layout-nlp-default / random*. Considering the small dataset, the public score can be sensitive to randomness even if you just change the random seed and leave all the other params unchanged. For more detailed discussion, you can refer to [this great forum](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436977).",
    "2517341": "Correct: the provided test data contains both public and private leaderboard data!"
  },
  "source": "meta"
}