{
  "id": 347202,
  "title": "Does the public/private split prone the shake-up ? ",
  "url": "/competitions/open-problems-multimodal/discussion/347202",
  "author_name": "",
  "post_date": "2022-08-23T09:38:05.501768Z",
  "votes": 16,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks for organizing that great competition !</p>\n<p>If I understand correctly private LB is quite different from public LB -<br>\non private we need to predict the unseen DAY (+1 donor), while on public - unseen DONOR  (but days are \"seen\"). <br>\nSee figure  in <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/data</a></p>\n<p>That seems kind of odd for me, and seems to be proning the shake up (if my understanding is correct?).<br>\nAny comments ? </p>\n<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
  "messages": [
    {
      "id": "1910255",
      "postDate": "08/23/2022 09:38:05",
      "content": "<p>Thanks for organizing that great competition !</p>\n<p>If I understand correctly private LB is quite different from public LB -<br>\non private we need to predict the unseen DAY (+1 donor), while on public - unseen DONOR  (but days are \"seen\"). <br>\nSee figure  in <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/data</a></p>\n<p>That seems kind of odd for me, and seems to be proning the shake up (if my understanding is correct?).<br>\nAny comments ? </p>\n<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
      "rawMarkdown": "Thanks for organizing that great competition !\n\nIf I understand correctly private LB is quite different from public LB -\non private we need to predict the unseen DAY (+1 donor), while on public - unseen DONOR  (but days are \"seen\"). \nSee figure  in https://www.kaggle.com/competitions/open-problems-multimodal/data\n\nThat seems kind of odd for me, and seems to be proning the shake up (if my understanding is correct?).\nAny comments ? \n\n@danielburkhardt",
      "votes": null
    },
    {
      "id": "1910916",
      "postDate": "08/23/2022 19:23:42",
      "content": "<p>Expression of some genes at day 7 seems to be quite different from the other days. <br>\nI am not saying that makes task unpredictable (cause other modalites might change in similar way - I need to look more on that), but at least it indicates caveats. I mean - split public/private LB seems not quite good.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F44f5f10436b63e2b2f2aad9fbdf76829%2Fphoto_2022-08-23_21-15-48.jpg?generation=1661282161756507&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Expression of some genes at day 7 seems to be quite different from the other days. \nI am not saying that makes task unpredictable (cause other modalites might change in similar way - I need to look more on that), but at least it indicates caveats. I mean - split public/private LB seems not quite good.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F44f5f10436b63e2b2f2aad9fbdf76829%2Fphoto_2022-08-23_21-15-48.jpg?generation=1661282161756507&alt=media)",
      "votes": null
    },
    {
      "id": "1911012",
      "postDate": "08/23/2022 21:37:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>! You are right in that there are significant differences between the train, private test, and public test data sets caused by the differences in donors and time points. In machine learning terms, this is called a <a href=\"https://en.wikipedia.org/wiki/Domain_adaptation#:~:text=A%20domain%20shift%2C%20or%20distributional,practical%20applications%20of%20artificial%20intelligence.\" target=\"_blank\">domain shift</a>, i.e. a change in the distribution between the train and the test data. However, this is intended - for several reasons:</p>\n<p>A. We want models to learn the underlying biological mechanisms - not just some arbitrary patterns of data. This is what is relevant in applications. A model that achieves this goal will be able to generalize better to a new donor or a new time point. <br>\nB. In biological applications, domain shifts are omnipresent: measurements that are performed on new donors, new laboratories, at slightly different times - in general other different experimental condition, which naturally leads to differences in the data sets.  A good model that can deployed to the real world is robust to such domain shifts. </p>\n<p>If you are concerned about how well your model can generalize to new time points or donors: the data sets allows you to evaluate that to an extent. For example, to test a model for generalization to a new time point for the Multiome data, train it on data from days 2,3,4 and evaluate it on predictions from day 7. That might give you a measure of how well your model can generalize. However, this is only a suggestion and there might be many ways of achieving good generalizability!</p>",
      "rawMarkdown": "Hi @alexandervc! You are right in that there are significant differences between the train, private test, and public test data sets caused by the differences in donors and time points. In machine learning terms, this is called a [domain shift](https://en.wikipedia.org/wiki/Domain_adaptation#:~:text=A%20domain%20shift%2C%20or%20distributional,practical%20applications%20of%20artificial%20intelligence.), i.e. a change in the distribution between the train and the test data. However, this is intended - for several reasons:\n\nA. We want models to learn the underlying biological mechanisms - not just some arbitrary patterns of data. This is what is relevant in applications. A model that achieves this goal will be able to generalize better to a new donor or a new time point. \nB. In biological applications, domain shifts are omnipresent: measurements that are performed on new donors, new laboratories, at slightly different times - in general other different experimental condition, which naturally leads to differences in the data sets.  A good model that can deployed to the real world is robust to such domain shifts. \n\nIf you are concerned about how well your model can generalize to new time points or donors: the data sets allows you to evaluate that to an extent. For example, to test a model for generalization to a new time point for the Multiome data, train it on data from days 2,3,4 and evaluate it on predictions from day 7. That might give you a measure of how well your model can generalize. However, this is only a suggestion and there might be many ways of achieving good generalizability!",
      "votes": null
    },
    {
      "id": "1913085",
      "postDate": "08/25/2022 06:08:16",
      "content": "<p><a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> <br>\nThank you for your answer ! <br>\nYes, I agree with all what you write, but from the research perspective )))<br>\nHere is kaggle - people compete for tiniest epsilon improvement of score, and such improvement is sensitive to tiny details  )))<br>\nI quite know about domain adaption problem, <br>\n(my colleagues recently wrote a paper and a package for it:<br>\n<a href=\"https://github.com/Mirkes/DAPCA\" target=\"_blank\">https://github.com/Mirkes/DAPCA</a> )<br>\nbut my concern is related , but  not precisely  that.<br>\nIn the language of the domain adaptation - my message is that domain adaption to public and to private LB are DIFFERENT,<br>\nand thus potentially leading to the shake up, and that is quite strange. <br>\nTypically people expect that public LB and private LB are more or less similar, at least not made deliberately different,<br>\notherwise what is the sense in public LB ?  </p>\n<p>In the other language - cross validation scheme good for private LB is different from the cross validation scheme good for public LB. </p>",
      "rawMarkdown": "peterholderrieth \nThank you for your answer ! \nYes, I agree with all what you write, but from the research perspective )))\nHere is kaggle - people compete for tiniest epsilon improvement of score, and such improvement is sensitive to tiny details  )))\nI quite know about domain adaption problem, \n(my colleagues recently wrote a paper and a package for it:\nhttps://github.com/Mirkes/DAPCA )\nbut my concern is related , but  not precisely  that.\nIn the language of the domain adaptation - my message is that domain adaption to public and to private LB are DIFFERENT,\nand thus potentially leading to the shake up, and that is quite strange. \nTypically people expect that public LB and private LB are more or less similar, at least not made deliberately different,\notherwise what is the sense in public LB ?  \n\nIn the other language - cross validation scheme good for private LB is different from the cross validation scheme good for public LB.",
      "votes": null
    },
    {
      "id": "1913674",
      "postDate": "08/25/2022 12:45:58",
      "content": "<p>Hi Alexander, you rightly note that the interest here in the public / private LB differences is in part a research question. For the public LB, we really want to see how well algorithms generalize across individuals. To win a prize, however, algorithms must generalize across time, which is the heart of the competition.</p>\n<p>I like Peter's suggestion to make use of the fact we've released 4 timepoints for 3 donors, and consider designing a custom CV scheme to help inform your own sense of how well your model generalizes to an unseen timepoint.</p>\n<p>I understand your concern that this may be not the most common setup for Kaggle. This is our first competition on the platform, and we made this decision in consultation with the data science team at Kaggle. Together we're keeping an eye on the public / private LB. We haven't seen anything outside of what one would expect in this context, but we'll continue to monitor.</p>",
      "rawMarkdown": "Hi Alexander, you rightly note that the interest here in the public / private LB differences is in part a research question. For the public LB, we really want to see how well algorithms generalize across individuals. To win a prize, however, algorithms must generalize across time, which is the heart of the competition.\n\nI like Peter's suggestion to make use of the fact we've released 4 timepoints for 3 donors, and consider designing a custom CV scheme to help inform your own sense of how well your model generalizes to an unseen timepoint.\n\nI understand your concern that this may be not the most common setup for Kaggle. This is our first competition on the platform, and we made this decision in consultation with the data science team at Kaggle. Together we're keeping an eye on the public / private LB. We haven't seen anything outside of what one would expect in this context, but we'll continue to monitor.",
      "votes": null
    },
    {
      "id": "1916272",
      "postDate": "08/27/2022 18:14:28",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nThank you for your answer !<br>\nThat is a great competition any way.<br>\nStill I would suggest to rearrange private/public LB split, such that they would be as identical as possible.<br>\nOtherwise it is not kind of fit the Kaggle spirit. </p>",
      "rawMarkdown": "danielburkhardt \nThank you for your answer !\nThat is a great competition any way.\nStill I would suggest to rearrange private/public LB split, such that they would be as identical as possible.\nOtherwise it is not kind of fit the Kaggle spirit.",
      "votes": null
    },
    {
      "id": "1916524",
      "postDate": "08/28/2022 00:40:54",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> ,</p>\n<p>Let me elaborate on the thought process a little bit. You're right that in the usual Kaggle setup what you want is to have the same relationship between Train / Public LB as Train / Private LB. With i.i.d. data, there's an easy solution to this problem with random sampling -- all you need to worry about is sizing your splits appropriately.</p>\n<p>With time dependent data (like we have in this competition), there's not always a great solution. We could put one day ahead in Public LB and two days ahead in Private LB, but dependence one day ahead isn't the same as dependence two days ahead, so this doesn't really solve the problem -- the Train / Public and Train / Private splits still don't have the same relationship. (And of course you can't put the same day in both Public and Private.)</p>\n<p>So instead we split a donor into the Public LB and gave you back the extra day. And as a competitor, which would you rather have? With the extra day, you can both train and validate up to a day before the Private LB day, but having the extra donor (and not the day) makes the problem harder (you have to predict two days ahead and have hardly any days to extrapolate from) without gaining you very much.</p>\n<p>So thanks for probing about this. It's good to understand these kinds of design decisions. I think there are probably good arguments for other kinds of designs, but we thought this design would go the furthest in answering the research question of interest.</p>",
      "rawMarkdown": "Hey @alexandervc ,\n\nLet me elaborate on the thought process a little bit. You're right that in the usual Kaggle setup what you want is to have the same relationship between Train / Public LB as Train / Private LB. With i.i.d. data, there's an easy solution to this problem with random sampling -- all you need to worry about is sizing your splits appropriately.\n\nWith time dependent data (like we have in this competition), there's not always a great solution. We could put one day ahead in Public LB and two days ahead in Private LB, but dependence one day ahead isn't the same as dependence two days ahead, so this doesn't really solve the problem -- the Train / Public and Train / Private splits still don't have the same relationship. (And of course you can't put the same day in both Public and Private.)\n\nSo instead we split a donor into the Public LB and gave you back the extra day. And as a competitor, which would you rather have? With the extra day, you can both train and validate up to a day before the Private LB day, but having the extra donor (and not the day) makes the problem harder (you have to predict two days ahead and have hardly any days to extrapolate from) without gaining you very much.\n\nSo thanks for probing about this. It's good to understand these kinds of design decisions. I think there are probably good arguments for other kinds of designs, but we thought this design would go the furthest in answering the research question of interest.",
      "votes": null
    },
    {
      "id": "1919049",
      "postDate": "08/30/2022 04:53:13",
      "content": "<p>Thank you for the comments!</p>\n<p>Why \"And of course you can't put the same day in both Public and Private\" ?</p>\n<p>Even for all conditions being the \"same\" , it is not easy to predict variation of one modality through another. (I mean/stress \"variation \" , not the averages, which are good baseline , but not what we want). Biologically e.g. on a way from rna to proteins , happens so many things which are unknown for modern science,  that it is clearly hard task.</p>\n<p>I might miss smth but discussed with others, no one understands. It is not timeseries task. We are not predicting future, but one modality through the other.</p>",
      "rawMarkdown": "Thank you for the comments!\n\nWhy \"And of course you can't put the same day in both Public and Private\" ?\n\nEven for all conditions being the \"same\" , it is not easy to predict variation of one modality through another. (I mean/stress \"variation \" , not the averages, which are good baseline , but not what we want). Biologically e.g. on a way from rna to proteins , happens so many things which are unknown for modern science,  that it is clearly hard task.\n\nI might miss smth but discussed with others, no one understands. It is not timeseries task. We are not predicting future, but one modality through the other.",
      "votes": null
    },
    {
      "id": "1919461",
      "postDate": "08/30/2022 12:36:21",
      "content": "<p>As Daniel mentioned, <code>algorithms must generalize across time, which is the heart of the competition</code>. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'</p>",
      "rawMarkdown": "As Daniel mentioned, `algorithms must generalize across time, which is the heart of the competition`. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'",
      "votes": null
    },
    {
      "id": "1925245",
      "postDate": "09/03/2022 18:55:38",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>  Hello, Ryan, thank you for your answer. <br>\nStill I cannot agree with it. Although, I can admit that you are right, because  it might be that there is something specific for that dataset,  such that it is indeed the case as you write. It might be it is some lesson from previous year competition. <br>\nIn any case I find that competition great, and problem with the validation is not that much important.</p>\n<p>But still. <br>\nI think the claim \"You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.' \"<br>\nMight be only justified by some specifics of that particular dataset, which you did not describe.<br>\nBut  <br>\nin standard situation  it is not true. <br>\nIf you overfit over public LB, that does not mean your result would be good for private LB. <br>\n<strong>Even if</strong> the input values of features are <strong>ARBITRARY CLOSE</strong> for public and private,<br>\nfor overfitted model the output values on private dataset can have <strong>ARBITRARY(!) GREAT ERROR</strong>.</p>\n<p><strong>Demonstration.</strong> To understand that (or mathematically prove)   - we should consider toy example of the linear regression.<br>\nSo imagine that you are solving the task to find the polynom p(x): p(x_i) = y_i<br>\nassume you allow youself polynomial of very high degree - to prone overfit.<br>\nAnd you overfitted the public LB - predicting the values for train+public LB - ABSOLUTELY CORRECTLY:<br>\np(x) = y - for all \"x\" in train+public.</p>\n<p>Does it imply that p(x) = y for \"x\" in private , even if \"x\" private is very very very close to some point  in train+public ? <br>\nNo. The answer is No. <br>\nAnd the error can be ARBITRARY GREAT (if you do not restrict the degree of \"p\" , in other words if you do not make regularization). <br>\nTo prove it we need to rely on basic mathematical properties of polynomials -  you can find \"p\": p(x_private) = 10^100 (any value), independently on what you have for p(train+priviate). (That can be seen from the Lagrange interpolation formula or in many other ways). </p>\n<p><strong>Thus the overfitting on public LB may lead to arbitrary big error on private LB.</strong> At least at that toy example and for many real life examples.</p>\n<p>Again, it might be that dataset is somewhat specific, but I do not see that has been spelled out. </p>",
      "rawMarkdown": "ryanholbrook  Hello, Ryan, thank you for your answer. \nStill I cannot agree with it. Although, I can admit that you are right, because  it might be that there is something specific for that dataset,  such that it is indeed the case as you write. It might be it is some lesson from previous year competition. \nIn any case I find that competition great, and problem with the validation is not that much important.\n\nBut still. \nI think the claim \"You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.' \"\nMight be only justified by some specifics of that particular dataset, which you did not describe.\nBut  \nin standard situation  it is not true. \nIf you overfit over public LB, that does not mean your result would be good for private LB. \n**Even if** the input values of features are **ARBITRARY CLOSE** for public and private,\nfor overfitted model the output values on private dataset can have **ARBITRARY(!) GREAT ERROR**.\n\n**Demonstration.** To understand that (or mathematically prove)   - we should consider toy example of the linear regression.\nSo imagine that you are solving the task to find the polynom p(x): p(x_i) = y_i\nassume you allow youself polynomial of very high degree - to prone overfit.\nAnd you overfitted the public LB - predicting the values for train+public LB - ABSOLUTELY CORRECTLY:\np(x) = y - for all \"x\" in train+public.\n\nDoes it imply that p(x) = y for \"x\" in private , even if \"x\" private is very very very close to some point  in train+public ? \nNo. The answer is No. \nAnd the error can be ARBITRARY GREAT (if you do not restrict the degree of \"p\" , in other words if you do not make regularization). \nTo prove it we need to rely on basic mathematical properties of polynomials -  you can find \"p\": p(x_private) = 10^100 (any value), independently on what you have for p(train+priviate). (That can be seen from the Lagrange interpolation formula or in many other ways). \n\n**Thus the overfitting on public LB may lead to arbitrary big error on private LB.** At least at that toy example and for many real life examples.\n\nAgain, it might be that dataset is somewhat specific, but I do not see that has been spelled out.",
      "votes": null
    },
    {
      "id": "1927487",
      "postDate": "09/05/2022 16:53:20",
      "content": "<blockquote>\n  <p>As Daniel mentioned, <code>algorithms must generalize across time, which is the heart of the competition</code>. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'</p>\n</blockquote>\n<p>This is really weird. Kaggle LB is traditionally designed to have public LB as an indicator of private LB. Given this big dataset, having day 7 in both public and private doesn't leak anything. It's just making the LB more trustful and the competition more interesting. You can't overfitting public LB without overfitting private LB if we don't know the indices of public/private. Hundreds of competitions had a LB this way. </p>",
      "rawMarkdown": "> As Daniel mentioned, `algorithms must generalize across time, which is the heart of the competition`. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'\n\nThis is really weird. Kaggle LB is traditionally designed to have public LB as an indicator of private LB. Given this big dataset, having day 7 in both public and private doesn't leak anything. It's just making the LB more trustful and the competition more interesting. You can't overfitting public LB without overfitting private LB if we don't know the indices of public/private. Hundreds of competitions had a LB this way.",
      "votes": null
    },
    {
      "id": "1927962",
      "postDate": "09/06/2022 05:56:01",
      "content": "<p>Thanks for the comment ! Exactly what I am trying to say. </p>",
      "rawMarkdown": "Thanks for the comment ! Exactly what I am trying to say.",
      "votes": null
    },
    {
      "id": "2028701",
      "postDate": "11/14/2022 07:21:53",
      "content": "<p>This is amazing, thanks a lot for sharing this</p>",
      "rawMarkdown": "This is amazing, thanks a lot for sharing this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1910916,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "08/23/2022 19:23:42",
      "content": "<p>Expression of some genes at day 7 seems to be quite different from the other days. <br>\nI am not saying that makes task unpredictable (cause other modalites might change in similar way - I need to look more on that), but at least it indicates caveats. I mean - split public/private LB seems not quite good.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F44f5f10436b63e2b2f2aad9fbdf76829%2Fphoto_2022-08-23_21-15-48.jpg?generation=1661282161756507&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1911012,
      "author_name": "peterholderrieth",
      "author_url": "",
      "post_date": "08/23/2022 21:37:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>! You are right in that there are significant differences between the train, private test, and public test data sets caused by the differences in donors and time points. In machine learning terms, this is called a <a href=\"https://en.wikipedia.org/wiki/Domain_adaptation#:~:text=A%20domain%20shift%2C%20or%20distributional,practical%20applications%20of%20artificial%20intelligence.\" target=\"_blank\">domain shift</a>, i.e. a change in the distribution between the train and the test data. However, this is intended - for several reasons:</p>\n<p>A. We want models to learn the underlying biological mechanisms - not just some arbitrary patterns of data. This is what is relevant in applications. A model that achieves this goal will be able to generalize better to a new donor or a new time point. <br>\nB. In biological applications, domain shifts are omnipresent: measurements that are performed on new donors, new laboratories, at slightly different times - in general other different experimental condition, which naturally leads to differences in the data sets.  A good model that can deployed to the real world is robust to such domain shifts. </p>\n<p>If you are concerned about how well your model can generalize to new time points or donors: the data sets allows you to evaluate that to an extent. For example, to test a model for generalization to a new time point for the Multiome data, train it on data from days 2,3,4 and evaluate it on predictions from day 7. That might give you a measure of how well your model can generalize. However, this is only a suggestion and there might be many ways of achieving good generalizability!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913085,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/25/2022 06:08:16",
          "content": "<p><a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> <br>\nThank you for your answer ! <br>\nYes, I agree with all what you write, but from the research perspective )))<br>\nHere is kaggle - people compete for tiniest epsilon improvement of score, and such improvement is sensitive to tiny details  )))<br>\nI quite know about domain adaption problem, <br>\n(my colleagues recently wrote a paper and a package for it:<br>\n<a href=\"https://github.com/Mirkes/DAPCA\" target=\"_blank\">https://github.com/Mirkes/DAPCA</a> )<br>\nbut my concern is related , but  not precisely  that.<br>\nIn the language of the domain adaptation - my message is that domain adaption to public and to private LB are DIFFERENT,<br>\nand thus potentially leading to the shake up, and that is quite strange. <br>\nTypically people expect that public LB and private LB are more or less similar, at least not made deliberately different,<br>\notherwise what is the sense in public LB ?  </p>\n<p>In the other language - cross validation scheme good for private LB is different from the cross validation scheme good for public LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1913674,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "08/25/2022 12:45:58",
          "content": "<p>Hi Alexander, you rightly note that the interest here in the public / private LB differences is in part a research question. For the public LB, we really want to see how well algorithms generalize across individuals. To win a prize, however, algorithms must generalize across time, which is the heart of the competition.</p>\n<p>I like Peter's suggestion to make use of the fact we've released 4 timepoints for 3 donors, and consider designing a custom CV scheme to help inform your own sense of how well your model generalizes to an unseen timepoint.</p>\n<p>I understand your concern that this may be not the most common setup for Kaggle. This is our first competition on the platform, and we made this decision in consultation with the data science team at Kaggle. Together we're keeping an eye on the public / private LB. We haven't seen anything outside of what one would expect in this context, but we'll continue to monitor.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1916272,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/27/2022 18:14:28",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nThank you for your answer !<br>\nThat is a great competition any way.<br>\nStill I would suggest to rearrange private/public LB split, such that they would be as identical as possible.<br>\nOtherwise it is not kind of fit the Kaggle spirit. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1916524,
          "author_name": "ryanholbrook",
          "author_url": "",
          "post_date": "08/28/2022 00:40:54",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> ,</p>\n<p>Let me elaborate on the thought process a little bit. You're right that in the usual Kaggle setup what you want is to have the same relationship between Train / Public LB as Train / Private LB. With i.i.d. data, there's an easy solution to this problem with random sampling -- all you need to worry about is sizing your splits appropriately.</p>\n<p>With time dependent data (like we have in this competition), there's not always a great solution. We could put one day ahead in Public LB and two days ahead in Private LB, but dependence one day ahead isn't the same as dependence two days ahead, so this doesn't really solve the problem -- the Train / Public and Train / Private splits still don't have the same relationship. (And of course you can't put the same day in both Public and Private.)</p>\n<p>So instead we split a donor into the Public LB and gave you back the extra day. And as a competitor, which would you rather have? With the extra day, you can both train and validate up to a day before the Private LB day, but having the extra donor (and not the day) makes the problem harder (you have to predict two days ahead and have hardly any days to extrapolate from) without gaining you very much.</p>\n<p>So thanks for probing about this. It's good to understand these kinds of design decisions. I think there are probably good arguments for other kinds of designs, but we thought this design would go the furthest in answering the research question of interest.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1919049,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/30/2022 04:53:13",
          "content": "<p>Thank you for the comments!</p>\n<p>Why \"And of course you can't put the same day in both Public and Private\" ?</p>\n<p>Even for all conditions being the \"same\" , it is not easy to predict variation of one modality through another. (I mean/stress \"variation \" , not the averages, which are good baseline , but not what we want). Biologically e.g. on a way from rna to proteins , happens so many things which are unknown for modern science,  that it is clearly hard task.</p>\n<p>I might miss smth but discussed with others, no one understands. It is not timeseries task. We are not predicting future, but one modality through the other.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1919461,
              "author_name": "ryanholbrook",
              "author_url": "",
              "post_date": "08/30/2022 12:36:21",
              "content": "<p>As Daniel mentioned, <code>algorithms must generalize across time, which is the heart of the competition</code>. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'</p>",
              "votes": null,
              "replies": [
                {
                  "id": 1927487,
                  "author_name": "khahuras",
                  "author_url": "",
                  "post_date": "09/05/2022 16:53:20",
                  "content": "<blockquote>\n  <p>As Daniel mentioned, <code>algorithms must generalize across time, which is the heart of the competition</code>. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'</p>\n</blockquote>\n<p>This is really weird. Kaggle LB is traditionally designed to have public LB as an indicator of private LB. Given this big dataset, having day 7 in both public and private doesn't leak anything. It's just making the LB more trustful and the competition more interesting. You can't overfitting public LB without overfitting private LB if we don't know the indices of public/private. Hundreds of competitions had a LB this way. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 1925245,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/03/2022 18:55:38",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>  Hello, Ryan, thank you for your answer. <br>\nStill I cannot agree with it. Although, I can admit that you are right, because  it might be that there is something specific for that dataset,  such that it is indeed the case as you write. It might be it is some lesson from previous year competition. <br>\nIn any case I find that competition great, and problem with the validation is not that much important.</p>\n<p>But still. <br>\nI think the claim \"You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.' \"<br>\nMight be only justified by some specifics of that particular dataset, which you did not describe.<br>\nBut  <br>\nin standard situation  it is not true. <br>\nIf you overfit over public LB, that does not mean your result would be good for private LB. <br>\n<strong>Even if</strong> the input values of features are <strong>ARBITRARY CLOSE</strong> for public and private,<br>\nfor overfitted model the output values on private dataset can have <strong>ARBITRARY(!) GREAT ERROR</strong>.</p>\n<p><strong>Demonstration.</strong> To understand that (or mathematically prove)   - we should consider toy example of the linear regression.<br>\nSo imagine that you are solving the task to find the polynom p(x): p(x_i) = y_i<br>\nassume you allow youself polynomial of very high degree - to prone overfit.<br>\nAnd you overfitted the public LB - predicting the values for train+public LB - ABSOLUTELY CORRECTLY:<br>\np(x) = y - for all \"x\" in train+public.</p>\n<p>Does it imply that p(x) = y for \"x\" in private , even if \"x\" private is very very very close to some point  in train+public ? <br>\nNo. The answer is No. <br>\nAnd the error can be ARBITRARY GREAT (if you do not restrict the degree of \"p\" , in other words if you do not make regularization). <br>\nTo prove it we need to rely on basic mathematical properties of polynomials -  you can find \"p\": p(x_private) = 10^100 (any value), independently on what you have for p(train+priviate). (That can be seen from the Lagrange interpolation formula or in many other ways). </p>\n<p><strong>Thus the overfitting on public LB may lead to arbitrary big error on private LB.</strong> At least at that toy example and for many real life examples.</p>\n<p>Again, it might be that dataset is somewhat specific, but I do not see that has been spelled out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1927962,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/06/2022 05:56:01",
          "content": "<p>Thanks for the comment ! Exactly what I am trying to say. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2028701,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2022 07:21:53",
      "content": "<p>This is amazing, thanks a lot for sharing this</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1910255": "Thanks for organizing that great competition !\n\nIf I understand correctly private LB is quite different from public LB -\non private we need to predict the unseen DAY (+1 donor), while on public - unseen DONOR  (but days are \"seen\"). \nSee figure  in https://www.kaggle.com/competitions/open-problems-multimodal/data\n\nThat seems kind of odd for me, and seems to be proning the shake up (if my understanding is correct?).\nAny comments ? \n\n@danielburkhardt",
    "1910916": "Expression of some genes at day 7 seems to be quite different from the other days. \nI am not saying that makes task unpredictable (cause other modalites might change in similar way - I need to look more on that), but at least it indicates caveats. I mean - split public/private LB seems not quite good.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F44f5f10436b63e2b2f2aad9fbdf76829%2Fphoto_2022-08-23_21-15-48.jpg?generation=1661282161756507&alt=media)",
    "1911012": "Hi @alexandervc! You are right in that there are significant differences between the train, private test, and public test data sets caused by the differences in donors and time points. In machine learning terms, this is called a [domain shift](https://en.wikipedia.org/wiki/Domain_adaptation#:~:text=A%20domain%20shift%2C%20or%20distributional,practical%20applications%20of%20artificial%20intelligence.), i.e. a change in the distribution between the train and the test data. However, this is intended - for several reasons:\n\nA. We want models to learn the underlying biological mechanisms - not just some arbitrary patterns of data. This is what is relevant in applications. A model that achieves this goal will be able to generalize better to a new donor or a new time point. \nB. In biological applications, domain shifts are omnipresent: measurements that are performed on new donors, new laboratories, at slightly different times - in general other different experimental condition, which naturally leads to differences in the data sets.  A good model that can deployed to the real world is robust to such domain shifts. \n\nIf you are concerned about how well your model can generalize to new time points or donors: the data sets allows you to evaluate that to an extent. For example, to test a model for generalization to a new time point for the Multiome data, train it on data from days 2,3,4 and evaluate it on predictions from day 7. That might give you a measure of how well your model can generalize. However, this is only a suggestion and there might be many ways of achieving good generalizability!",
    "1913085": "peterholderrieth \nThank you for your answer ! \nYes, I agree with all what you write, but from the research perspective )))\nHere is kaggle - people compete for tiniest epsilon improvement of score, and such improvement is sensitive to tiny details  )))\nI quite know about domain adaption problem, \n(my colleagues recently wrote a paper and a package for it:\nhttps://github.com/Mirkes/DAPCA )\nbut my concern is related , but  not precisely  that.\nIn the language of the domain adaptation - my message is that domain adaption to public and to private LB are DIFFERENT,\nand thus potentially leading to the shake up, and that is quite strange. \nTypically people expect that public LB and private LB are more or less similar, at least not made deliberately different,\notherwise what is the sense in public LB ?  \n\nIn the other language - cross validation scheme good for private LB is different from the cross validation scheme good for public LB.",
    "1913674": "Hi Alexander, you rightly note that the interest here in the public / private LB differences is in part a research question. For the public LB, we really want to see how well algorithms generalize across individuals. To win a prize, however, algorithms must generalize across time, which is the heart of the competition.\n\nI like Peter's suggestion to make use of the fact we've released 4 timepoints for 3 donors, and consider designing a custom CV scheme to help inform your own sense of how well your model generalizes to an unseen timepoint.\n\nI understand your concern that this may be not the most common setup for Kaggle. This is our first competition on the platform, and we made this decision in consultation with the data science team at Kaggle. Together we're keeping an eye on the public / private LB. We haven't seen anything outside of what one would expect in this context, but we'll continue to monitor.",
    "1916272": "danielburkhardt \nThank you for your answer !\nThat is a great competition any way.\nStill I would suggest to rearrange private/public LB split, such that they would be as identical as possible.\nOtherwise it is not kind of fit the Kaggle spirit.",
    "1916524": "Hey @alexandervc ,\n\nLet me elaborate on the thought process a little bit. You're right that in the usual Kaggle setup what you want is to have the same relationship between Train / Public LB as Train / Private LB. With i.i.d. data, there's an easy solution to this problem with random sampling -- all you need to worry about is sizing your splits appropriately.\n\nWith time dependent data (like we have in this competition), there's not always a great solution. We could put one day ahead in Public LB and two days ahead in Private LB, but dependence one day ahead isn't the same as dependence two days ahead, so this doesn't really solve the problem -- the Train / Public and Train / Private splits still don't have the same relationship. (And of course you can't put the same day in both Public and Private.)\n\nSo instead we split a donor into the Public LB and gave you back the extra day. And as a competitor, which would you rather have? With the extra day, you can both train and validate up to a day before the Private LB day, but having the extra donor (and not the day) makes the problem harder (you have to predict two days ahead and have hardly any days to extrapolate from) without gaining you very much.\n\nSo thanks for probing about this. It's good to understand these kinds of design decisions. I think there are probably good arguments for other kinds of designs, but we thought this design would go the furthest in answering the research question of interest.",
    "1919049": "Thank you for the comments!\n\nWhy \"And of course you can't put the same day in both Public and Private\" ?\n\nEven for all conditions being the \"same\" , it is not easy to predict variation of one modality through another. (I mean/stress \"variation \" , not the averages, which are good baseline , but not what we want). Biologically e.g. on a way from rna to proteins , happens so many things which are unknown for modern science,  that it is clearly hard task.\n\nI might miss smth but discussed with others, no one understands. It is not timeseries task. We are not predicting future, but one modality through the other.",
    "1919461": "As Daniel mentioned, `algorithms must generalize across time, which is the heart of the competition`. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'",
    "1925245": "ryanholbrook  Hello, Ryan, thank you for your answer. \nStill I cannot agree with it. Although, I can admit that you are right, because  it might be that there is something specific for that dataset,  such that it is indeed the case as you write. It might be it is some lesson from previous year competition. \nIn any case I find that competition great, and problem with the validation is not that much important.\n\nBut still. \nI think the claim \"You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.' \"\nMight be only justified by some specifics of that particular dataset, which you did not describe.\nBut  \nin standard situation  it is not true. \nIf you overfit over public LB, that does not mean your result would be good for private LB. \n**Even if** the input values of features are **ARBITRARY CLOSE** for public and private,\nfor overfitted model the output values on private dataset can have **ARBITRARY(!) GREAT ERROR**.\n\n**Demonstration.** To understand that (or mathematically prove)   - we should consider toy example of the linear regression.\nSo imagine that you are solving the task to find the polynom p(x): p(x_i) = y_i\nassume you allow youself polynomial of very high degree - to prone overfit.\nAnd you overfitted the public LB - predicting the values for train+public LB - ABSOLUTELY CORRECTLY:\np(x) = y - for all \"x\" in train+public.\n\nDoes it imply that p(x) = y for \"x\" in private , even if \"x\" private is very very very close to some point  in train+public ? \nNo. The answer is No. \nAnd the error can be ARBITRARY GREAT (if you do not restrict the degree of \"p\" , in other words if you do not make regularization). \nTo prove it we need to rely on basic mathematical properties of polynomials -  you can find \"p\": p(x_private) = 10^100 (any value), independently on what you have for p(train+priviate). (That can be seen from the Lagrange interpolation formula or in many other ways). \n\n**Thus the overfitting on public LB may lead to arbitrary big error on private LB.** At least at that toy example and for many real life examples.\n\nAgain, it might be that dataset is somewhat specific, but I do not see that has been spelled out.",
    "1927487": "> As Daniel mentioned, `algorithms must generalize across time, which is the heart of the competition`. So for CITEseq, we couldn't have day 7 in both the Public and Private sets, since the public set then leaks information about the private set. You could make your score unrealistically good by overfitting to the public LB, and then your private LB score isn't a reliable measure of 'generalization across time.'\n\nThis is really weird. Kaggle LB is traditionally designed to have public LB as an indicator of private LB. Given this big dataset, having day 7 in both public and private doesn't leak anything. It's just making the LB more trustful and the competition more interesting. You can't overfitting public LB without overfitting private LB if we don't know the indices of public/private. Hundreds of competitions had a LB this way.",
    "1927962": "Thanks for the comment ! Exactly what I am trying to say.",
    "2028701": "This is amazing, thanks a lot for sharing this"
  },
  "source": "meta"
}