{
  "id": 461159,
  "title": "20th Place Solution Writeup For Open Problems - Single-cell Perturbations Competition",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/jalil-nourisa-antoine-passemie-20th-place-solution",
  "author_name": "",
  "post_date": "2023-12-18T07:39:24.853Z",
  "votes": 11,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Please find attached our detailed solution write-up, submitted for consideration for the Judge's reward. For your convenience, we have included all necessary citations within the attached PDF document.</p>\n<p>Additionally, the individual notebooks referenced in our write-up are accessible at the following GitHub repository: <a href=\"https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master\" target=\"_blank\">https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master</a></p>",
  "messages": [
    {
      "id": "2559525",
      "postDate": "12/12/2023 22:43:10",
      "content": "<p>Please find attached our detailed solution write-up, submitted for consideration for the Judge's reward. For your convenience, we have included all necessary citations within the attached PDF document.</p>\n<p>Additionally, the individual notebooks referenced in our write-up are accessible at the following GitHub repository: <a href=\"https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master\" target=\"_blank\">https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master</a></p>",
      "rawMarkdown": "Please find attached our detailed solution write-up, submitted for consideration for the Judge's reward. For your convenience, we have included all necessary citations within the attached PDF document.\n\nAdditionally, the individual notebooks referenced in our write-up are accessible at the following GitHub repository: [https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master](https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master)",
      "votes": null
    },
    {
      "id": "2565253",
      "postDate": "12/17/2023 20:31:27",
      "content": "<p>Great analysis ! Congratulations with the medal ! <br>\nWould it be possible to share the notebooks for your analysis ? <br>\nPS<br>\nConcerning Figure 2 <br>\nNot fully understand the \"x-axis\" - it is ln( p-value) ? but ln(0.001) is -6.90, on the plot - we see something around -18 </p>\n<p>PS<br>\nTo complement your Figure 1 and illustrate uniform distribution for p-values (under null hypothisis)<br>\n for simply-minded folks , that may be useful:<br>\n(Left jump-up seems to present biological signal)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdac7a56a2545c2021b7da692e7d78d43%2Fphoto_2023-12-17_22-18-38.jpg?generation=1702848044668884&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155423176&amp;cellId=6\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155423176&amp;cellId=6</a></p>",
      "rawMarkdown": "Great analysis ! Congratulations with the medal ! \nWould it be possible to share the notebooks for your analysis ? \nPS\nConcerning Figure 2 \nNot fully understand the \"x-axis\" - it is ln( p-value) ? but ln(0.001) is -6.90, on the plot - we see something around -18 \n\nPS\nTo complement your Figure 1 and illustrate uniform distribution for p-values (under null hypothisis)\n for simply-minded folks , that may be useful:\n(Left jump-up seems to present biological signal)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdac7a56a2545c2021b7da692e7d78d43%2Fphoto_2023-12-17_22-18-38.jpg?generation=1702848044668884&alt=media)\nhttps://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155423176&cellId=6",
      "votes": null
    },
    {
      "id": "2565676",
      "postDate": "12/18/2023 08:09:47",
      "content": "<p>thank you for the comment. the code for this section is here <a href=\"https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb\" target=\"_blank\">https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb</a>. </p>\n<p>np.log(.001/18206)=-16.7, as we correct for multiple testing. </p>",
      "rawMarkdown": "thank you for the comment. the code for this section is here [https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb](https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb). \n\nnp.log(.001/18206)=-16.7, as we correct for multiple testing.",
      "votes": null
    },
    {
      "id": "2566748",
      "postDate": "12/19/2023 05:58:21",
      "content": "<p>Thank you very much for the reply  !</p>\n<p>PS<br>\nTo complement your Figure 2 (great finding on housekeeping genes, appreciate it a lot) , <br>\none may also (may be also would be helpful for simply-minded folks):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F1c7e78e463830df1d1bc84ce98ca44ee%2FScreenshot%202023-12-19%20065354.png?generation=1702965348401417&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11</a></p>\n<p>PSPS<br>\nA small misprint page 15: \"Chirvov [13]\" -&gt; \"ChErvov [13]\"    :)))</p>",
      "rawMarkdown": "Thank you very much for the reply  !\n\nPS\nTo complement your Figure 2 (great finding on housekeeping genes, appreciate it a lot) , \none may also (may be also would be helpful for simply-minded folks):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F1c7e78e463830df1d1bc84ce98ca44ee%2FScreenshot%202023-12-19%20065354.png?generation=1702965348401417&alt=media)\nhttps://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&cellId=11\n\nPSPS\nA small misprint page 15: \"Chirvov [13]\" -> \"ChErvov [13]\"    :)))",
      "votes": null
    },
    {
      "id": "2566803",
      "postDate": "12/19/2023 07:03:02",
      "content": "<p>thank you very much for the comment, and my bad in mistyping your name! i will certainly fix it. best </p>",
      "rawMarkdown": "thank you very much for the comment, and my bad in mistyping your name! i will certainly fix it. best",
      "votes": null
    },
    {
      "id": "2566816",
      "postDate": "12/19/2023 07:11:00",
      "content": "<p>Fantastic work <a href=\"https://www.kaggle.com/jalilnourisa\" target=\"_blank\">@jalilnourisa</a> !</p>",
      "rawMarkdown": "Fantastic work @jalilnourisa !",
      "votes": null
    },
    {
      "id": "2566822",
      "postDate": "12/19/2023 07:16:22",
      "content": "<p>thank you </p>",
      "rawMarkdown": "thank you",
      "votes": null
    },
    {
      "id": "2568355",
      "postDate": "12/20/2023 12:59:33",
      "content": "<p>One more comment/question:<br>\nDo not you think that these exceptional samples on Figure 3,4,5 of yours  - may be the same \"bad\" samples which appear in <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> analysis. Which can also be characterized as those where  percent of strongly perturbed genes is exceptionally big (my complement to AmbrosM post - reproduced below here). Also these sample can be characterized as the biggest error for the prediction models - see Antonina Dolgorukova analysis: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663</a> </p>\n<p>Let me reproduce - first - your Figures 3,4 and  later  from my post</p>\n<p>By the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F760e2c153f2200aeafa2351a47c5a9a5%2FUntitled.png?generation=1703077808306458&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F53639fd8935c8cdfdc031a555f004fc8%2FScreenshot%202023-12-20%20145348.png?generation=1703080635376231&amp;alt=media\" alt=\"\"></p>\n<p>(Post: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661#2566894\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661#2566894</a> )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F30c4bd14eecb278ce6949ebdff3ca819%2Finbox_2262596_4e1c24530cbda8ef941825e5c464268c_Screenshot%202023-12-19%20092655.png?generation=1703076985506023&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdee832b1ab7a13829f75bdddef9eadde%2Finbox_2262596_4e84fbb28aade383843e4a7ec4cdcf66_Screenshot%202023-12-19%20100930.png?generation=1703076996730520&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "One more comment/question:\nDo not you think that these exceptional samples on Figure 3,4,5 of yours  - may be the same \"bad\" samples which appear in @ambrosm analysis. Which can also be characterized as those where  percent of strongly perturbed genes is exceptionally big (my complement to AmbrosM post - reproduced below here). Also these sample can be characterized as the biggest error for the prediction models - see Antonina Dolgorukova analysis: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663 \n\n\nLet me reproduce - first - your Figures 3,4 and  later  from my post\n\nBy the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F760e2c153f2200aeafa2351a47c5a9a5%2FUntitled.png?generation=1703077808306458&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F53639fd8935c8cdfdc031a555f004fc8%2FScreenshot%202023-12-20%20145348.png?generation=1703080635376231&alt=media)\n\n(Post: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661#2566894 )\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F30c4bd14eecb278ce6949ebdff3ca819%2Finbox_2262596_4e1c24530cbda8ef941825e5c464268c_Screenshot%202023-12-19%20092655.png?generation=1703076985506023&alt=media)\n\nhttps://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&cellId=10\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdee832b1ab7a13829f75bdddef9eadde%2Finbox_2262596_4e84fbb28aade383843e4a7ec4cdcf66_Screenshot%202023-12-19%20100930.png?generation=1703076996730520&alt=media)",
      "votes": null
    },
    {
      "id": "2568459",
      "postDate": "12/20/2023 14:52:36",
      "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> for your deeper look.</p>\n<p>These different analyses all pointing to the same direction. Thanks for sharing your point of view.</p>\n<p>Re \"By the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)\" Its MLN 2238 (the color is confusing, i agree). </p>",
      "rawMarkdown": "Thank you very much, @alexandervc for your deeper look.\n\nThese different analyses all pointing to the same direction. Thanks for sharing your point of view.\n\nRe \"By the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)\" Its MLN 2238 (the color is confusing, i agree).",
      "votes": null
    },
    {
      "id": "2572048",
      "postDate": "12/23/2023 20:25:43",
      "content": "<p>Your analysis of MRRMSE pitfalls is quite striking - in particular presented on Figure 7  ! (screenshotted below here).</p>\n<p>And related to the questions which bother me a lot: to what extent we are predicting biological effects, rather than various artefacts of Limma/whatever ? What can be good demonstrations that we are predicting biology ? (@malteluecken , <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>  )  What are the good metrics to measure biological part of prediction ? And what pre-processing is better ?  </p>\n<p>Let me rephrase your analysis as follows - our targets contain - a) small number of big values - bio-signal b) large number of small values - noise (uniform at 0-1 for p-values). The question is:  what dominates the MRRMSE - (a) or (b), i.e. bio-signal or noise  ? </p>\n<p>Your analysis shows: that MRRMSE is dominated by small values (below say 2, i.e. p-value =0.01) - which are mostly noise, but not biological signal. (These values correspond to p-values above 0.01 - which means not a signal, but just random uniformly distributed noise which arise when null hypothesis is true. Uniform distribution is clearly seen).  </p>\n<p>So it might seem that it strongly forbids use of MRRSME as a bio-meaningful metric, however I am not completely sure about that, because of the following:</p>\n<p>On the one hand - indeed -  your analysis demonstrates that MRRMSE is bad metric for local training the models, because it is dominated by noisy (not bio meaningful) values and thus forces models to learn the noise-part which is completely meaningless - noise is random - it will not generalize to unseen data. </p>\n<p>However, I am not so sure how bad it is for evaluating the models performance on the unseen data, because: most probably noise part of values is unpredictable and all models would have more or less random predictions on that part - and so law of large numbers makes all models in a sense equal on that part noisy part, and what makes one prediction better than the other - is prediction of  big biologically meaningful values.  Not sure that is correct argument. Would be happy to hear your opinion. </p>\n<p>PS </p>\n<p>May be  we should think how to split genes in \"pure noise\" and \"bio signal\" , choose different strategies to predict the two classes.</p>\n<p>Any way, let me complement your analysis by the following:<br>\nLet me consider your  MRRMSE-corrected  metrics with different thresholds: 1,2,3 and  explore how well that LOCAL metric is correlated with the public LB MRRMSE-score. Outcomes - mostly confirm your insight:   <strong>MRRMSE-corrected is BETTER correlated with LB, than standard MRRMSE</strong>. That is especially for <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> CV-scheme, also for random scheme, strangely enough for MT-scheme is NOT the case. See the screenshot: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team</a></p>\n<p>(The analysis might be biased - by choice of the models, on the other hand, my collection of models is quite diverse - but that was done before Pyboost and other strong models.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F729d1e81e3365bcb3752b7d8d2653e70%2FScreenshot%202023-12-23%20201950.png?generation=1703359959463139&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F240a8cd3eca13fe1d88dce76d3999b8f%2Fmrrmse_analysis.png?generation=1703365889728981&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fd15a1a0b595507d5edfbbe3d089218cb%2FScreenshot%202023-12-23%20210635.png?generation=1703363177358159&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Your analysis of MRRMSE pitfalls is quite striking - in particular presented on Figure 7  ! (screenshotted below here).\n\nAnd related to the questions which bother me a lot: to what extent we are predicting biological effects, rather than various artefacts of Limma/whatever ? What can be good demonstrations that we are predicting biology ? (@malteluecken , @danielburkhardt  )  What are the good metrics to measure biological part of prediction ? And what pre-processing is better ?  \n\nLet me rephrase your analysis as follows - our targets contain - a) small number of big values - bio-signal b) large number of small values - noise (uniform at 0-1 for p-values). The question is:  what dominates the MRRMSE - (a) or (b), i.e. bio-signal or noise  ? \n\nYour analysis shows: that MRRMSE is dominated by small values (below say 2, i.e. p-value =0.01) - which are mostly noise, but not biological signal. (These values correspond to p-values above 0.01 - which means not a signal, but just random uniformly distributed noise which arise when null hypothesis is true. Uniform distribution is clearly seen).  \n\nSo it might seem that it strongly forbids use of MRRSME as a bio-meaningful metric, however I am not completely sure about that, because of the following:\n\nOn the one hand - indeed -  your analysis demonstrates that MRRMSE is bad metric for local training the models, because it is dominated by noisy (not bio meaningful) values and thus forces models to learn the noise-part which is completely meaningless - noise is random - it will not generalize to unseen data. \n\nHowever, I am not so sure how bad it is for evaluating the models performance on the unseen data, because: most probably noise part of values is unpredictable and all models would have more or less random predictions on that part - and so law of large numbers makes all models in a sense equal on that part noisy part, and what makes one prediction better than the other - is prediction of  big biologically meaningful values.  Not sure that is correct argument. Would be happy to hear your opinion. \n\nPS \n\nMay be  we should think how to split genes in \"pure noise\" and \"bio signal\" , choose different strategies to predict the two classes.\n\n\nAny way, let me complement your analysis by the following:\nLet me consider your  MRRMSE-corrected  metrics with different thresholds: 1,2,3 and  explore how well that LOCAL metric is correlated with the public LB MRRMSE-score. Outcomes - mostly confirm your insight:   **MRRMSE-corrected is BETTER correlated with LB, than standard MRRMSE**. That is especially for @ambrosm CV-scheme, also for random scheme, strangely enough for MT-scheme is NOT the case. See the screenshot: \nhttps://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\n\n(The analysis might be biased - by choice of the models, on the other hand, my collection of models is quite diverse - but that was done before Pyboost and other strong models.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F729d1e81e3365bcb3752b7d8d2653e70%2FScreenshot%202023-12-23%20201950.png?generation=1703359959463139&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F240a8cd3eca13fe1d88dce76d3999b8f%2Fmrrmse_analysis.png?generation=1703365889728981&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fd15a1a0b595507d5edfbbe3d089218cb%2FScreenshot%202023-12-23%20210635.png?generation=1703363177358159&alt=media)",
      "votes": null
    },
    {
      "id": "2572603",
      "postDate": "12/24/2023 11:40:44",
      "content": "<p>Thanks for the feedback <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>, and thank you for the follow-up investigations you made. The numbers you show are quite intriguing and promising, as they show that our revised version of MRRMSE could be even more relevant for internal CV than we originally thought. It also suggests that I might have been too stringent by using 3 as a threshold, and that a threshold of 1 might be more suitable. But let's remain cautious, because the CV-LB relationship is probably too complex to draw any conclusion here. I would have expected the correlation between CV and LB to increase if <em>both</em> were based on the corrected version of MRRMSE.</p>\n<p>But let me venture an explanation of why the LB-CV correlation is lower for the corrected MRRMSE when using the MT scheme. With the MT scheme, the positive controls (Dabrafenib and Belinostat) will end up in the validation set much more frequently than in other CV schemes, and these controls carry a lot of signal (many high DE values). Because of this, the CV should in principle penalize conservative models more, and encourage predictions that deviate from the null distribution. However, the test set does not contain positive controls, and therefore contains less significant values. Consequence is that they will penalize the less conservative models a bit more, hence the difference between CV and LB. This probably does not explain everything, but this is my best bet at the moment. Also, your analysis is based on only 47 data points (which I guess took a lot of efforts to collect), which makes me wonder how reliable these correlation coefficients are. Did you compute p-values or confidence intervals?</p>\n<p>Regarding your suggestion that MRRMSE is probably not as bad in practice thanks to the law of large numbers, I think the intuition makes sense since the sum of noise terms would converge to the same value and therefore appear as a simple \"bias term\" in the MRRMSE that would be the same for all models. However, this implicitly assumes that all models behave the same on the non-DE genes (for example, their predictions all follow a Gaussian distribution with same location and scale), and it is definitely not the case. For most models, the mode of predictions (the \"peak\" in the distribution) is strictly greater than 0, and its exact value depends on the model you choose. Also, the variance of the predictions depends on whether or not you used ensembling, and how many models you ensembled.</p>",
      "rawMarkdown": "Thanks for the feedback @alexandervc, and thank you for the follow-up investigations you made. The numbers you show are quite intriguing and promising, as they show that our revised version of MRRMSE could be even more relevant for internal CV than we originally thought. It also suggests that I might have been too stringent by using 3 as a threshold, and that a threshold of 1 might be more suitable. But let's remain cautious, because the CV-LB relationship is probably too complex to draw any conclusion here. I would have expected the correlation between CV and LB to increase if *both* were based on the corrected version of MRRMSE.\n\nBut let me venture an explanation of why the LB-CV correlation is lower for the corrected MRRMSE when using the MT scheme. With the MT scheme, the positive controls (Dabrafenib and Belinostat) will end up in the validation set much more frequently than in other CV schemes, and these controls carry a lot of signal (many high DE values). Because of this, the CV should in principle penalize conservative models more, and encourage predictions that deviate from the null distribution. However, the test set does not contain positive controls, and therefore contains less significant values. Consequence is that they will penalize the less conservative models a bit more, hence the difference between CV and LB. This probably does not explain everything, but this is my best bet at the moment. Also, your analysis is based on only 47 data points (which I guess took a lot of efforts to collect), which makes me wonder how reliable these correlation coefficients are. Did you compute p-values or confidence intervals?\n\nRegarding your suggestion that MRRMSE is probably not as bad in practice thanks to the law of large numbers, I think the intuition makes sense since the sum of noise terms would converge to the same value and therefore appear as a simple \"bias term\" in the MRRMSE that would be the same for all models. However, this implicitly assumes that all models behave the same on the non-DE genes (for example, their predictions all follow a Gaussian distribution with same location and scale), and it is definitely not the case. For most models, the mode of predictions (the \"peak\" in the distribution) is strictly greater than 0, and its exact value depends on the model you choose. Also, the variance of the predictions depends on whether or not you used ensembling, and how many models you ensembled.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2565253,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/17/2023 20:31:27",
      "content": "<p>Great analysis ! Congratulations with the medal ! <br>\nWould it be possible to share the notebooks for your analysis ? <br>\nPS<br>\nConcerning Figure 2 <br>\nNot fully understand the \"x-axis\" - it is ln( p-value) ? but ln(0.001) is -6.90, on the plot - we see something around -18 </p>\n<p>PS<br>\nTo complement your Figure 1 and illustrate uniform distribution for p-values (under null hypothisis)<br>\n for simply-minded folks , that may be useful:<br>\n(Left jump-up seems to present biological signal)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdac7a56a2545c2021b7da692e7d78d43%2Fphoto_2023-12-17_22-18-38.jpg?generation=1702848044668884&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155423176&amp;cellId=6\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155423176&amp;cellId=6</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2565676,
          "author_name": "jalilnourisa",
          "author_url": "",
          "post_date": "12/18/2023 08:09:47",
          "content": "<p>thank you for the comment. the code for this section is here <a href=\"https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb\" target=\"_blank\">https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb</a>. </p>\n<p>np.log(.001/18206)=-16.7, as we correct for multiple testing. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2566748,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "12/19/2023 05:58:21",
              "content": "<p>Thank you very much for the reply  !</p>\n<p>PS<br>\nTo complement your Figure 2 (great finding on housekeeping genes, appreciate it a lot) , <br>\none may also (may be also would be helpful for simply-minded folks):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F1c7e78e463830df1d1bc84ce98ca44ee%2FScreenshot%202023-12-19%20065354.png?generation=1702965348401417&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11</a></p>\n<p>PSPS<br>\nA small misprint page 15: \"Chirvov [13]\" -&gt; \"ChErvov [13]\"    :)))</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2566803,
                  "author_name": "jalilnourisa",
                  "author_url": "",
                  "post_date": "12/19/2023 07:03:02",
                  "content": "<p>thank you very much for the comment, and my bad in mistyping your name! i will certainly fix it. best </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2566816,
      "author_name": "arviinndn",
      "author_url": "",
      "post_date": "12/19/2023 07:11:00",
      "content": "<p>Fantastic work <a href=\"https://www.kaggle.com/jalilnourisa\" target=\"_blank\">@jalilnourisa</a> !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2566822,
          "author_name": "jalilnourisa",
          "author_url": "",
          "post_date": "12/19/2023 07:16:22",
          "content": "<p>thank you </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2568355,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/20/2023 12:59:33",
      "content": "<p>One more comment/question:<br>\nDo not you think that these exceptional samples on Figure 3,4,5 of yours  - may be the same \"bad\" samples which appear in <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> analysis. Which can also be characterized as those where  percent of strongly perturbed genes is exceptionally big (my complement to AmbrosM post - reproduced below here). Also these sample can be characterized as the biggest error for the prediction models - see Antonina Dolgorukova analysis: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663</a> </p>\n<p>Let me reproduce - first - your Figures 3,4 and  later  from my post</p>\n<p>By the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F760e2c153f2200aeafa2351a47c5a9a5%2FUntitled.png?generation=1703077808306458&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F53639fd8935c8cdfdc031a555f004fc8%2FScreenshot%202023-12-20%20145348.png?generation=1703080635376231&amp;alt=media\" alt=\"\"></p>\n<p>(Post: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661#2566894\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661#2566894</a> )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F30c4bd14eecb278ce6949ebdff3ca819%2Finbox_2262596_4e1c24530cbda8ef941825e5c464268c_Screenshot%202023-12-19%20092655.png?generation=1703076985506023&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdee832b1ab7a13829f75bdddef9eadde%2Finbox_2262596_4e84fbb28aade383843e4a7ec4cdcf66_Screenshot%202023-12-19%20100930.png?generation=1703076996730520&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2568459,
          "author_name": "jalilnourisa",
          "author_url": "",
          "post_date": "12/20/2023 14:52:36",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> for your deeper look.</p>\n<p>These different analyses all pointing to the same direction. Thanks for sharing your point of view.</p>\n<p>Re \"By the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)\" Its MLN 2238 (the color is confusing, i agree). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2572048,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/23/2023 20:25:43",
      "content": "<p>Your analysis of MRRMSE pitfalls is quite striking - in particular presented on Figure 7  ! (screenshotted below here).</p>\n<p>And related to the questions which bother me a lot: to what extent we are predicting biological effects, rather than various artefacts of Limma/whatever ? What can be good demonstrations that we are predicting biology ? (@malteluecken , <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>  )  What are the good metrics to measure biological part of prediction ? And what pre-processing is better ?  </p>\n<p>Let me rephrase your analysis as follows - our targets contain - a) small number of big values - bio-signal b) large number of small values - noise (uniform at 0-1 for p-values). The question is:  what dominates the MRRMSE - (a) or (b), i.e. bio-signal or noise  ? </p>\n<p>Your analysis shows: that MRRMSE is dominated by small values (below say 2, i.e. p-value =0.01) - which are mostly noise, but not biological signal. (These values correspond to p-values above 0.01 - which means not a signal, but just random uniformly distributed noise which arise when null hypothesis is true. Uniform distribution is clearly seen).  </p>\n<p>So it might seem that it strongly forbids use of MRRSME as a bio-meaningful metric, however I am not completely sure about that, because of the following:</p>\n<p>On the one hand - indeed -  your analysis demonstrates that MRRMSE is bad metric for local training the models, because it is dominated by noisy (not bio meaningful) values and thus forces models to learn the noise-part which is completely meaningless - noise is random - it will not generalize to unseen data. </p>\n<p>However, I am not so sure how bad it is for evaluating the models performance on the unseen data, because: most probably noise part of values is unpredictable and all models would have more or less random predictions on that part - and so law of large numbers makes all models in a sense equal on that part noisy part, and what makes one prediction better than the other - is prediction of  big biologically meaningful values.  Not sure that is correct argument. Would be happy to hear your opinion. </p>\n<p>PS </p>\n<p>May be  we should think how to split genes in \"pure noise\" and \"bio signal\" , choose different strategies to predict the two classes.</p>\n<p>Any way, let me complement your analysis by the following:<br>\nLet me consider your  MRRMSE-corrected  metrics with different thresholds: 1,2,3 and  explore how well that LOCAL metric is correlated with the public LB MRRMSE-score. Outcomes - mostly confirm your insight:   <strong>MRRMSE-corrected is BETTER correlated with LB, than standard MRRMSE</strong>. That is especially for <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> CV-scheme, also for random scheme, strangely enough for MT-scheme is NOT the case. See the screenshot: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team</a></p>\n<p>(The analysis might be biased - by choice of the models, on the other hand, my collection of models is quite diverse - but that was done before Pyboost and other strong models.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F729d1e81e3365bcb3752b7d8d2653e70%2FScreenshot%202023-12-23%20201950.png?generation=1703359959463139&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F240a8cd3eca13fe1d88dce76d3999b8f%2Fmrrmse_analysis.png?generation=1703365889728981&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fd15a1a0b595507d5edfbbe3d089218cb%2FScreenshot%202023-12-23%20210635.png?generation=1703363177358159&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2572603,
          "author_name": "antoinepassemiers",
          "author_url": "",
          "post_date": "12/24/2023 11:40:44",
          "content": "<p>Thanks for the feedback <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>, and thank you for the follow-up investigations you made. The numbers you show are quite intriguing and promising, as they show that our revised version of MRRMSE could be even more relevant for internal CV than we originally thought. It also suggests that I might have been too stringent by using 3 as a threshold, and that a threshold of 1 might be more suitable. But let's remain cautious, because the CV-LB relationship is probably too complex to draw any conclusion here. I would have expected the correlation between CV and LB to increase if <em>both</em> were based on the corrected version of MRRMSE.</p>\n<p>But let me venture an explanation of why the LB-CV correlation is lower for the corrected MRRMSE when using the MT scheme. With the MT scheme, the positive controls (Dabrafenib and Belinostat) will end up in the validation set much more frequently than in other CV schemes, and these controls carry a lot of signal (many high DE values). Because of this, the CV should in principle penalize conservative models more, and encourage predictions that deviate from the null distribution. However, the test set does not contain positive controls, and therefore contains less significant values. Consequence is that they will penalize the less conservative models a bit more, hence the difference between CV and LB. This probably does not explain everything, but this is my best bet at the moment. Also, your analysis is based on only 47 data points (which I guess took a lot of efforts to collect), which makes me wonder how reliable these correlation coefficients are. Did you compute p-values or confidence intervals?</p>\n<p>Regarding your suggestion that MRRMSE is probably not as bad in practice thanks to the law of large numbers, I think the intuition makes sense since the sum of noise terms would converge to the same value and therefore appear as a simple \"bias term\" in the MRRMSE that would be the same for all models. However, this implicitly assumes that all models behave the same on the non-DE genes (for example, their predictions all follow a Gaussian distribution with same location and scale), and it is definitely not the case. For most models, the mode of predictions (the \"peak\" in the distribution) is strictly greater than 0, and its exact value depends on the model you choose. Also, the variance of the predictions depends on whether or not you used ensembling, and how many models you ensembled.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2559525": "Please find attached our detailed solution write-up, submitted for consideration for the Judge's reward. For your convenience, we have included all necessary citations within the attached PDF document.\n\nAdditionally, the individual notebooks referenced in our write-up are accessible at the following GitHub repository: [https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master](https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/tree/master)",
    "2565253": "Great analysis ! Congratulations with the medal ! \nWould it be possible to share the notebooks for your analysis ? \nPS\nConcerning Figure 2 \nNot fully understand the \"x-axis\" - it is ln( p-value) ? but ln(0.001) is -6.90, on the plot - we see something around -18 \n\nPS\nTo complement your Figure 1 and illustrate uniform distribution for p-values (under null hypothisis)\n for simply-minded folks , that may be useful:\n(Left jump-up seems to present biological signal)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdac7a56a2545c2021b7da692e7d78d43%2Fphoto_2023-12-17_22-18-38.jpg?generation=1702848044668884&alt=media)\nhttps://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155423176&cellId=6",
    "2565676": "thank you for the comment. the code for this section is here [https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb](https://github.com/AntoinePassemiers/Open-Challenges-Single-Cell-Perturbations/blob/master/op2-de-dl.ipynb). \n\nnp.log(.001/18206)=-16.7, as we correct for multiple testing.",
    "2566748": "Thank you very much for the reply  !\n\nPS\nTo complement your Figure 2 (great finding on housekeeping genes, appreciate it a lot) , \none may also (may be also would be helpful for simply-minded folks):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F1c7e78e463830df1d1bc84ce98ca44ee%2FScreenshot%202023-12-19%20065354.png?generation=1702965348401417&alt=media)\nhttps://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&cellId=11\n\nPSPS\nA small misprint page 15: \"Chirvov [13]\" -> \"ChErvov [13]\"    :)))",
    "2566803": "thank you very much for the comment, and my bad in mistyping your name! i will certainly fix it. best",
    "2566816": "Fantastic work @jalilnourisa !",
    "2566822": "thank you",
    "2568355": "One more comment/question:\nDo not you think that these exceptional samples on Figure 3,4,5 of yours  - may be the same \"bad\" samples which appear in @ambrosm analysis. Which can also be characterized as those where  percent of strongly perturbed genes is exceptionally big (my complement to AmbrosM post - reproduced below here). Also these sample can be characterized as the biggest error for the prediction models - see Antonina Dolgorukova analysis: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663 \n\n\nLet me reproduce - first - your Figures 3,4 and  later  from my post\n\nBy the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F760e2c153f2200aeafa2351a47c5a9a5%2FUntitled.png?generation=1703077808306458&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F53639fd8935c8cdfdc031a555f004fc8%2FScreenshot%202023-12-20%20145348.png?generation=1703080635376231&alt=media)\n\n(Post: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661#2566894 )\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F30c4bd14eecb278ce6949ebdff3ca819%2Finbox_2262596_4e1c24530cbda8ef941825e5c464268c_Screenshot%202023-12-19%20092655.png?generation=1703076985506023&alt=media)\n\nhttps://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&cellId=10\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fdee832b1ab7a13829f75bdddef9eadde%2Finbox_2262596_4e84fbb28aade383843e4a7ec4cdcf66_Screenshot%202023-12-19%20100930.png?generation=1703076996730520&alt=media)",
    "2568459": "Thank you very much, @alexandervc for your deeper look.\n\nThese different analyses all pointing to the same direction. Thanks for sharing your point of view.\n\nRe \"By the way those highlighted is it MLN2238 or Porcn Inhibitor ? (Color seems to be the same)\" Its MLN 2238 (the color is confusing, i agree).",
    "2572048": "Your analysis of MRRMSE pitfalls is quite striking - in particular presented on Figure 7  ! (screenshotted below here).\n\nAnd related to the questions which bother me a lot: to what extent we are predicting biological effects, rather than various artefacts of Limma/whatever ? What can be good demonstrations that we are predicting biology ? (@malteluecken , @danielburkhardt  )  What are the good metrics to measure biological part of prediction ? And what pre-processing is better ?  \n\nLet me rephrase your analysis as follows - our targets contain - a) small number of big values - bio-signal b) large number of small values - noise (uniform at 0-1 for p-values). The question is:  what dominates the MRRMSE - (a) or (b), i.e. bio-signal or noise  ? \n\nYour analysis shows: that MRRMSE is dominated by small values (below say 2, i.e. p-value =0.01) - which are mostly noise, but not biological signal. (These values correspond to p-values above 0.01 - which means not a signal, but just random uniformly distributed noise which arise when null hypothesis is true. Uniform distribution is clearly seen).  \n\nSo it might seem that it strongly forbids use of MRRSME as a bio-meaningful metric, however I am not completely sure about that, because of the following:\n\nOn the one hand - indeed -  your analysis demonstrates that MRRMSE is bad metric for local training the models, because it is dominated by noisy (not bio meaningful) values and thus forces models to learn the noise-part which is completely meaningless - noise is random - it will not generalize to unseen data. \n\nHowever, I am not so sure how bad it is for evaluating the models performance on the unseen data, because: most probably noise part of values is unpredictable and all models would have more or less random predictions on that part - and so law of large numbers makes all models in a sense equal on that part noisy part, and what makes one prediction better than the other - is prediction of  big biologically meaningful values.  Not sure that is correct argument. Would be happy to hear your opinion. \n\nPS \n\nMay be  we should think how to split genes in \"pure noise\" and \"bio signal\" , choose different strategies to predict the two classes.\n\n\nAny way, let me complement your analysis by the following:\nLet me consider your  MRRMSE-corrected  metrics with different thresholds: 1,2,3 and  explore how well that LOCAL metric is correlated with the public LB MRRMSE-score. Outcomes - mostly confirm your insight:   **MRRMSE-corrected is BETTER correlated with LB, than standard MRRMSE**. That is especially for @ambrosm CV-scheme, also for random scheme, strangely enough for MT-scheme is NOT the case. See the screenshot: \nhttps://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\n\n(The analysis might be biased - by choice of the models, on the other hand, my collection of models is quite diverse - but that was done before Pyboost and other strong models.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F729d1e81e3365bcb3752b7d8d2653e70%2FScreenshot%202023-12-23%20201950.png?generation=1703359959463139&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F240a8cd3eca13fe1d88dce76d3999b8f%2Fmrrmse_analysis.png?generation=1703365889728981&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fd15a1a0b595507d5edfbbe3d089218cb%2FScreenshot%202023-12-23%20210635.png?generation=1703363177358159&alt=media)",
    "2572603": "Thanks for the feedback @alexandervc, and thank you for the follow-up investigations you made. The numbers you show are quite intriguing and promising, as they show that our revised version of MRRMSE could be even more relevant for internal CV than we originally thought. It also suggests that I might have been too stringent by using 3 as a threshold, and that a threshold of 1 might be more suitable. But let's remain cautious, because the CV-LB relationship is probably too complex to draw any conclusion here. I would have expected the correlation between CV and LB to increase if *both* were based on the corrected version of MRRMSE.\n\nBut let me venture an explanation of why the LB-CV correlation is lower for the corrected MRRMSE when using the MT scheme. With the MT scheme, the positive controls (Dabrafenib and Belinostat) will end up in the validation set much more frequently than in other CV schemes, and these controls carry a lot of signal (many high DE values). Because of this, the CV should in principle penalize conservative models more, and encourage predictions that deviate from the null distribution. However, the test set does not contain positive controls, and therefore contains less significant values. Consequence is that they will penalize the less conservative models a bit more, hence the difference between CV and LB. This probably does not explain everything, but this is my best bet at the moment. Also, your analysis is based on only 47 data points (which I guess took a lot of efforts to collect), which makes me wonder how reliable these correlation coefficients are. Did you compute p-values or confidence intervals?\n\nRegarding your suggestion that MRRMSE is probably not as bad in practice thanks to the law of large numbers, I think the intuition makes sense since the sum of noise terms would converge to the same value and therefore appear as a simple \"bias term\" in the MRRMSE that would be the same for all models. However, this implicitly assumes that all models behave the same on the non-DE genes (for example, their predictions all follow a Gaussian distribution with same location and scale), and it is definitely not the case. For most models, the mode of predictions (the \"peak\" in the distribution) is strictly greater than 0, and its exact value depends on the model you choose. Also, the variance of the predictions depends on whether or not you used ensembling, and how many models you ensembled."
  },
  "source": "meta"
}