{
  "id": 73572,
  "title": "Gaussian Process Intro",
  "url": "/competitions/PLAsTiCC-2018/discussion/73572",
  "author_name": "CPMP",
  "post_date": "2018-12-04T05:34:50.322000",
  "votes": 34,
  "comment_count": 34,
  "views": 0,
  "content": "<p>If you are like me, i.e. willing to give a try at gaussian process modeling, but having no clue about it, then you may want to read this series of introductory posts: <a href=\"https://allofyourbases.com/category/gaussian-process-modelling/\">https://allofyourbases.com/category/gaussian-process-modelling/</a></p>\n\n<p>The bible is this book, available online: <a href=\"http://www.gaussianprocess.org/gpml/chapters/\">http://www.gaussianprocess.org/gpml/chapters/</a></p>\n\n<p>There is also this paper: <a href=\"http://www.robots.ox.ac.uk/~sjrob/Pubs/philTransA_2012.pdf\">http://www.robots.ox.ac.uk/~sjrob/Pubs/philTransA_2012.pdf</a></p>\n\n<p>Good read!</p>",
  "messages": [
    {
      "id": 432624,
      "postDate": "2018-12-04T05:34:50.323Z",
      "content": "<p>If you are like me, i.e. willing to give a try at gaussian process modeling, but having no clue about it, then you may want to read this series of introductory posts: <a href=\"https://allofyourbases.com/category/gaussian-process-modelling/\">https://allofyourbases.com/category/gaussian-process-modelling/</a></p>\n\n<p>The bible is this book, available online: <a href=\"http://www.gaussianprocess.org/gpml/chapters/\">http://www.gaussianprocess.org/gpml/chapters/</a></p>\n\n<p>There is also this paper: <a href=\"http://www.robots.ox.ac.uk/~sjrob/Pubs/philTransA_2012.pdf\">http://www.robots.ox.ac.uk/~sjrob/Pubs/philTransA_2012.pdf</a></p>\n\n<p>Good read!</p>",
      "rawMarkdown": "If you are like me, i.e. willing to give a try at gaussian process modeling, but having no clue about it, then you may want to read this series of introductory posts: https://allofyourbases.com/category/gaussian-process-modelling/\n\nThe bible is this book, available online: http://www.gaussianprocess.org/gpml/chapters/\n\nThere is also this paper: http://www.robots.ox.ac.uk/~sjrob/Pubs/philTransA_2012.pdf\n\nGood read!",
      "votes": 34
    },
    {
      "id": 438397,
      "postDate": "2018-12-13T15:38:06.247Z",
      "content": "<p>Some nice gp fit\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/438397/10873/gp.png\" alt=\"gp fit\"></p>\n\n<p>This also shows the limit of gaussian process modeling.  The peak was clearly before the second observation period, hence the extrapolation is a bit shaky.  But the curves look smoother that the autoencoder ones that have been shared so far.</p>",
      "rawMarkdown": "Some nice gp fit\n![gp fit][1]\n\nThis also shows the limit of gaussian process modeling.  The peak was clearly before the second observation period, hence the extrapolation is a bit shaky.  But the curves look smoother that the autoencoder ones that have been shared so far.\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/438397/10873/gp.png",
      "votes": 6,
      "replies": [
        {
          "id": 438555,
          "postDate": "2018-12-13T21:57:34.740Z",
          "content": "<p>I hope you will share your gp code after the competition! :)</p>",
          "rawMarkdown": "I hope you will share your gp code after the competition! :)",
          "votes": 2
        },
        {
          "id": 438583,
          "postDate": "2018-12-13T23:04:12.973Z",
          "content": "<p>I will!</p>",
          "rawMarkdown": "I will!",
          "votes": 2
        }
      ]
    },
    {
      "id": 435135,
      "postDate": "2018-12-07T15:11:17.173Z",
      "content": "<p>From my own tests, GP is slow because you have inside a matrix inversion. Then you run it multiple time when optmizing the hyperparameters.\nI tested instead Kernel methods for time interpolation, and suprisingly achieved better Likelihood with the same number of parameters. (spread of the data around 0 when no input is present and time window length). Fitting only for the time length and setting the sigma to the std of the time serie  reachs almost as good likelihood, while running even faster.</p>\n\n<p>I agree with Mithrillon that having a template to extrapolate / interpolate lightcurve would help. However the exact template depends on the class... This is the research I'm doing now. For an easy class, you can look at my kernel for class 6:\n<a href=\"https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6/comments\">https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6/comments</a> </p>",
      "rawMarkdown": "From my own tests, GP is slow because you have inside a matrix inversion. Then you run it multiple time when optmizing the hyperparameters.\nI tested instead Kernel methods for time interpolation, and suprisingly achieved better Likelihood with the same number of parameters. (spread of the data around 0 when no input is present and time window length). Fitting only for the time length and setting the sigma to the std of the time serie  reachs almost as good likelihood, while running even faster.\n\nI agree with Mithrillon that having a template to extrapolate / interpolate lightcurve would help. However the exact template depends on the class... This is the research I'm doing now. For an easy class, you can look at my kernel for class 6:\nhttps://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6/comments ",
      "votes": 4
    },
    {
      "id": 433020,
      "postDate": "2018-12-04T15:12:37.167Z",
      "content": "<p>Is there anyone find GP useful?</p>",
      "rawMarkdown": "Is there anyone find GP useful?",
      "votes": 1,
      "replies": [
        {
          "id": 433923,
          "postDate": "2018-12-05T16:52:20.380Z",
          "content": "<p>I hope it is ;)</p>",
          "rawMarkdown": "I hope it is ;)"
        },
        {
          "id": 434230,
          "postDate": "2018-12-06T04:37:13.553Z",
          "content": "<p>Not taking part in this competition, but Gaussian Processes were my jam for 8-10 months in machine learning.  They can be used for almost everything from ensembling to beer making.  Just taking a look at this competition, seems GPs may be useful.  However, for this competition especially -- don't let your optimization processes get stuck in local minima and remember that CV scores are always going to be more reliable than GPs.</p>",
          "rawMarkdown": "Not taking part in this competition, but Gaussian Processes were my jam for 8-10 months in machine learning.  They can be used for almost everything from ensembling to beer making.  Just taking a look at this competition, seems GPs may be useful.  However, for this competition especially -- don't let your optimization processes get stuck in local minima and remember that CV scores are always going to be more reliable than GPs."
        },
        {
          "id": 434238,
          "postDate": "2018-12-06T04:53:29.680Z",
          "content": "<blockquote>\n  <p>CV scores are always going to be more reliable than GPs.</p>\n</blockquote>\n\n<p>You are new here, hence you may not have seen yet that in this competition CV scores are not reliable because test data is significantly different from train data.</p>\n\n<p>But my real reason for replying is that I don't understand how you can compare a CV score with a GP.  A GP is not a score, it is?</p>",
          "rawMarkdown": "&gt; CV scores are always going to be more reliable than GPs.\n\nYou are new here, hence you may not have seen yet that in this competition CV scores are not reliable because test data is significantly different from train data.\n\nBut my real reason for replying is that I don't understand how you can compare a CV score with a GP.  A GP is not a score, it is?",
          "votes": 1
        },
        {
          "id": 434249,
          "postDate": "2018-12-06T05:23:43.390Z",
          "content": "<p>Oh, didn't see that CV scores are unreliable.</p>\n\n<p>And I was referring to a different use of GPs than I think you may have originally intended.  Don't know about their usefulness here; probably too expensive to train on the dataset and depending on class balance (not sure here what that is like) could have a bad balance of exploration/exploitation.</p>\n\n<p>Anyways to respond to your actual question: GPs are not a score but they can be used to find the minimum/maximum of an unknown function (such as the score of an ml model given hyperparameters) very quickly.  It may be tempting to assume that the GP has identified the minimum/maximum score for the ml model and what hyperparameters give that score, but you need to fall back on your CV scores to ultimately determine what hyperparameters you are using.  My personal experience with this is that I chose my model hyperparameters for final submission based on what the GP had as earning me the maximum score; however, it would have been much better to the model that had the best local CV (I dropped like 42% of the entire leaderboard because of this mistake).</p>\n\n<p>Sorry for the ambiguity.  You are totally right that CV scores and GPs can't be directly compared.</p>\n\n<p>You're right: I'm new here.  Looking forward to learning more from people who are way smarter than me.</p>",
          "rawMarkdown": "Oh, didn't see that CV scores are unreliable.\n\nAnd I was referring to a different use of GPs than I think you may have originally intended.  Don't know about their usefulness here; probably too expensive to train on the dataset and depending on class balance (not sure here what that is like) could have a bad balance of exploration/exploitation.\n\nAnyways to respond to your actual question: GPs are not a score but they can be used to find the minimum/maximum of an unknown function (such as the score of an ml model given hyperparameters) very quickly.  It may be tempting to assume that the GP has identified the minimum/maximum score for the ml model and what hyperparameters give that score, but you need to fall back on your CV scores to ultimately determine what hyperparameters you are using.  My personal experience with this is that I chose my model hyperparameters for final submission based on what the GP had as earning me the maximum score; however, it would have been much better to the model that had the best local CV (I dropped like 42% of the entire leaderboard because of this mistake).\n\nSorry for the ambiguity.  You are totally right that CV scores and GPs can't be directly compared.\n\nYou're right: I'm new here.  Looking forward to learning more from people who are way smarter than me.",
          "votes": 1
        },
        {
          "id": 434263,
          "postDate": "2018-12-06T05:55:32.687Z",
          "content": "<p>OK, you mean using GP to optimize hyperparameters.  This is a very legitimate use case.</p>",
          "rawMarkdown": "OK, you mean using GP to optimize hyperparameters.  This is a very legitimate use case.",
          "votes": 1
        }
      ]
    },
    {
      "id": 432731,
      "postDate": "2018-12-04T09:02:17.643Z",
      "content": "<p>I have tried this one but it takes too long to fit just the training set so i dropped the idea</p>",
      "rawMarkdown": "I have tried this one but it takes too long to fit just the training set so i dropped the idea",
      "votes": 1,
      "replies": [
        {
          "id": 432753,
          "postDate": "2018-12-04T09:29:26.543Z",
          "content": "<p>Yeah, GP is pretty slow (typically n^3). I think the other flaw is that it does not model the unseen portion of the data very well, so you can expect all sorts of weirdness being fitted if you have a supernova with only the rising phase. It is used with success in many previous researches, but from what I have read, they all had much more complete light curves to work with. I think a good curve model should at least be moderately resistant to partially obscuring the light curve, whereas samples drawn from a GP are usually completely dependent on the observed points for a given object, so it might not fare well with the particualr dataset we have.</p>",
          "rawMarkdown": "Yeah, GP is pretty slow (typically n^3). I think the other flaw is that it does not model the unseen portion of the data very well, so you can expect all sorts of weirdness being fitted if you have a supernova with only the rising phase. It is used with success in many previous researches, but from what I have read, they all had much more complete light curves to work with. I think a good curve model should at least be moderately resistant to partially obscuring the light curve, whereas samples drawn from a GP are usually completely dependent on the observed points for a given object, so it might not fare well with the particualr dataset we have.",
          "votes": 1
        },
        {
          "id": 432885,
          "postDate": "2018-12-04T12:36:34.363Z",
          "content": "<blockquote>\n  <p>it does not model the unseen portion of the data very well</p>\n</blockquote>\n\n<p>Do you think autoencoders model unseen part better?</p>",
          "rawMarkdown": "&gt; it does not model the unseen portion of the data very well\n\nDo you think autoencoders model unseen part better?"
        },
        {
          "id": 432901,
          "postDate": "2018-12-04T12:53:42.557Z",
          "content": "<p>Not necessarily so. It'd be much easier if we had \"complete\" light curves as reference, in which case it becomes a curve restoration task. Piecing together broken light curves is much harder, and I still don't have a great solution for that. I can get an autoencoder to default to an \"average\" light curve rather than the mean with insufficient data, but extrapolation is still questionable. At least I was able to mostly eliminate instant drop to zero or straight line slope down to the next value now.</p>",
          "rawMarkdown": "Not necessarily so. It'd be much easier if we had \"complete\" light curves as reference, in which case it becomes a curve restoration task. Piecing together broken light curves is much harder, and I still don't have a great solution for that. I can get an autoencoder to default to an \"average\" light curve rather than the mean with insufficient data, but extrapolation is still questionable. At least I was able to mostly eliminate instant drop to zero or straight line slope down to the next value now."
        }
      ]
    },
    {
      "id": 435907,
      "postDate": "2018-12-09T03:31:09.060Z",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>Seriously, how do you manage to learn all this so fast? You wrote <code>having no clue about it</code> 5 days ago and now finishing test run ... I feel like damn blonde now...</p>",
      "rawMarkdown": "@cpmpml\n\nSeriously, how do you manage to learn all this so fast? You wrote ```having no clue about it``` 5 days ago and now finishing test run ... I feel like damn blonde now...",
      "votes": 2,
      "replies": [
        {
          "id": 435950,
          "postDate": "2018-12-09T06:11:23.580Z",
          "content": "<p>This is funny, I always think I am a slow learner.  There is so much I don't know, and so much new things produced constantly, especially in deep learning.</p>\n\n<p>To compensate I look for relevant material and read it carefully.  For instance, how many of us here noticed the comment about celerite being faster than other gp packages in what we got from organizers?</p>\n\n<p>Back to gp, sure, I can compute stuff, but making good use of it is another story.</p>",
          "rawMarkdown": "This is funny, I always think I am a slow learner.  There is so much I don't know, and so much new things produced constantly, especially in deep learning.\n\nTo compensate I look for relevant material and read it carefully.  For instance, how many of us here noticed the comment about celerite being faster than other gp packages in what we got from organizers?\n\nBack to gp, sure, I can compute stuff, but making good use of it is another story.",
          "votes": 3
        }
      ]
    },
    {
      "id": 433911,
      "postDate": "2018-12-05T16:37:27.927Z",
      "content": "<p>Looks pretty to me.  But it is not that pretty in many cases, still experimenting.  And being pretty does not mean useful either.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/433911/10832/pg.png\" alt=\"gaussian process\"></p>",
      "rawMarkdown": "Looks pretty to me.  But it is not that pretty in many cases, still experimenting.  And being pretty does not mean useful either.\n\n![gaussian process][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/433911/10832/pg.png",
      "votes": 2,
      "replies": [
        {
          "id": 434021,
          "postDate": "2018-12-05T19:34:40.207Z",
          "content": "<p>Thanks for this. How much time does it take to for you to compute this for the training set?</p>",
          "rawMarkdown": "Thanks for this. How much time does it take to for you to compute this for the training set?"
        },
        {
          "id": 434074,
          "postDate": "2018-12-05T21:11:29.457Z",
          "content": "<p>Hey CPMP - alot of the measurements look like they were done over 3 time periods - just saying if you can partition the data you will really speed up your Gaussian Processing</p>",
          "rawMarkdown": "Hey CPMP - alot of the measurements look like they were done over 3 time periods - just saying if you can partition the data you will really speed up your Gaussian Processing"
        },
        {
          "id": 434125,
          "postDate": "2018-12-05T23:31:02.817Z",
          "content": "<p>But then you don't get the benefit of interpolating unseen values...</p>",
          "rawMarkdown": "But then you don't get the benefit of interpolating unseen values...",
          "votes": 1
        },
        {
          "id": 434236,
          "postDate": "2018-12-06T04:50:41.173Z",
          "content": "<blockquote>\n  <p>How much time does it take to for you to compute this for the training set?</p>\n</blockquote>\n\n<p>About 1h30 with a single thread.  it means it can be computed in an affordable time on test data with some parallelism.</p>\n\n<blockquote>\n  <p>if you can partition the data you will really speed up your Gaussian Processing</p>\n</blockquote>\n\n<p>As @Mithrillion wrote, this defeats the purpose.</p>",
          "rawMarkdown": "&gt; How much time does it take to for you to compute this for the training set?\n\nAbout 1h30 with a single thread.  it means it can be computed in an affordable time on test data with some parallelism.\n\n&gt;  if you can partition the data you will really speed up your Gaussian Processing\n\nAs @Mithrillion wrote, this defeats the purpose.",
          "votes": 1
        },
        {
          "id": 434632,
          "postDate": "2018-12-06T17:48:12.567Z",
          "content": "<p>This example looks like very good data example, not all objects are like that. </p>\n\n<p>BTW About 1h30 with a single thread -- it's on CPU, correct? Out of curiosity, how long do you think it will run for the test ? I've already paralleled periods on 24 CPUs and had to drop the idea after few days of running... </p>",
          "rawMarkdown": "This example looks like very good data example, not all objects are like that. \n\nBTW About 1h30 with a single thread -- it's on CPU, correct? Out of curiosity, how long do you think it will run for the test ? I've already paralleled periods on 24 CPUs and had to drop the idea after few days of running... "
        },
        {
          "id": 434641,
          "postDate": "2018-12-06T18:06:50.797Z",
          "content": "<p>Using 20 threads it should take a coupe of days.  I am still experimenting on how to best use the posterior of the gp.  </p>",
          "rawMarkdown": "Using 20 threads it should take a coupe of days.  I am still experimenting on how to best use the posterior of the gp.  "
        },
        {
          "id": 435754,
          "postDate": "2018-12-08T17:34:20.780Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> 3 days passed, did you try it on test, did it help your model ?</p>",
          "rawMarkdown": "@cpmpml 3 days passed, did you try it on test, did it help your model ?"
        },
        {
          "id": 435758,
          "postDate": "2018-12-08T17:58:19.060Z",
          "content": "<p>I am computing on test data, will see result tomorrow I think.</p>",
          "rawMarkdown": "I am computing on test data, will see result tomorrow I think."
        },
        {
          "id": 435905,
          "postDate": "2018-12-09T03:27:33.810Z",
          "content": "<p>Ok, I'll write down your current LB :-)</p>",
          "rawMarkdown": "Ok, I'll write down your current LB :-)"
        },
        {
          "id": 436220,
          "postDate": "2018-12-09T23:38:50.107Z",
          "content": "<p>LB did not change... so probably making use of GP is still problematic... is it ? We still need to find out how to get another 0.1 gain, time is shrinking...</p>",
          "rawMarkdown": "LB did not change... so probably making use of GP is still problematic... is it ? We still need to find out how to get another 0.1 gain, time is shrinking..."
        },
        {
          "id": 436638,
          "postDate": "2018-12-10T17:09:47.423Z",
          "content": "<p>We're busy trying stuff.  Ensembling can wait a bit...</p>",
          "rawMarkdown": "We're busy trying stuff.  Ensembling can wait a bit..."
        },
        {
          "id": 436659,
          "postDate": "2018-12-10T17:56:23.627Z",
          "content": "<p>So did it help on single model ? (I know you don't have to answer) :)  We are busy running instances, so trying to plan computational resources.... next competition I choose with the small test size! <br>\n(My God, I got addicted already, when it's not too late to quit :-)))?) </p>",
          "rawMarkdown": "So did it help on single model ? (I know you don't have to answer) :)  We are busy running instances, so trying to plan computational resources.... next competition I choose with the small test size!  \n(My God, I got addicted already, when it's not too late to quit :-)))?) ",
          "votes": 2
        },
        {
          "id": 436668,
          "postDate": "2018-12-10T18:05:40.207Z",
          "content": "<p>Wait until the Kaggle Blues kick in once the competition finishes! ;)</p>",
          "rawMarkdown": "Wait until the Kaggle Blues kick in once the competition finishes! ;)",
          "votes": 1
        },
        {
          "id": 436670,
          "postDate": "2018-12-10T18:12:20.840Z",
          "content": "<p>@Blonde, I shared a lot already.  I am fine sharing pointer to publicly available information, but I won't tell you what works and what does not work.  Why don't you ask the leaders?  ;)</p>",
          "rawMarkdown": "@Blonde, I shared a lot already.  I am fine sharing pointer to publicly available information, but I won't tell you what works and what does not work.  Why don't you ask the leaders?  ;)",
          "votes": -1
        },
        {
          "id": 436672,
          "postDate": "2018-12-10T18:14:38.033Z",
          "content": "<p><a href=\"/scirpus\">@scirpus</a> </p>\n\n<p>I am actually on maternity, my Kaggle Blues is treating my Babies Blues... very effective treatment I must say :)</p>",
          "rawMarkdown": "@scirpus \n\nI am actually on maternity, my Kaggle Blues is treating my Babies Blues... very effective treatment I must say :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 435542,
      "postDate": "2018-12-08T07:30:56.700Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 435661,
          "postDate": "2018-12-08T13:44:30.873Z",
          "content": "<p>I thought I recommended a textbook in my post...</p>",
          "rawMarkdown": "I thought I recommended a textbook in my post..."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 438397,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-13T15:38:06.247000",
      "content": "<p>Some nice gp fit\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/438397/10873/gp.png\" alt=\"gp fit\"></p>\n\n<p>This also shows the limit of gaussian process modeling.  The peak was clearly before the second observation period, hence the extrapolation is a bit shaky.  But the curves look smoother that the autoencoder ones that have been shared so far.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 438555,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2018-12-13T21:57:34.740000",
          "content": "<p>I hope you will share your gp code after the competition! :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 438583,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-13T23:04:12.973000",
          "content": "<p>I will!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 435135,
      "author_name": "Manu Gangler",
      "author_url": "",
      "post_date": "2018-12-07T15:11:17.173000",
      "content": "<p>From my own tests, GP is slow because you have inside a matrix inversion. Then you run it multiple time when optmizing the hyperparameters.\nI tested instead Kernel methods for time interpolation, and suprisingly achieved better Likelihood with the same number of parameters. (spread of the data around 0 when no input is present and time window length). Fitting only for the time length and setting the sigma to the std of the time serie  reachs almost as good likelihood, while running even faster.</p>\n\n<p>I agree with Mithrillon that having a template to extrapolate / interpolate lightcurve would help. However the exact template depends on the class... This is the research I'm doing now. For an easy class, you can look at my kernel for class 6:\n<a href=\"https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6/comments\">https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6/comments</a> </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 433020,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "2018-12-04T15:12:37.167000",
      "content": "<p>Is there anyone find GP useful?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433923,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-05T16:52:20.380000",
          "content": "<p>I hope it is ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434230,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-12-06T04:37:13.553000",
          "content": "<p>Not taking part in this competition, but Gaussian Processes were my jam for 8-10 months in machine learning.  They can be used for almost everything from ensembling to beer making.  Just taking a look at this competition, seems GPs may be useful.  However, for this competition especially -- don't let your optimization processes get stuck in local minima and remember that CV scores are always going to be more reliable than GPs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434238,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-06T04:53:29.680000",
          "content": "<blockquote>\n  <p>CV scores are always going to be more reliable than GPs.</p>\n</blockquote>\n\n<p>You are new here, hence you may not have seen yet that in this competition CV scores are not reliable because test data is significantly different from train data.</p>\n\n<p>But my real reason for replying is that I don't understand how you can compare a CV score with a GP.  A GP is not a score, it is?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434249,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-12-06T05:23:43.390000",
          "content": "<p>Oh, didn't see that CV scores are unreliable.</p>\n\n<p>And I was referring to a different use of GPs than I think you may have originally intended.  Don't know about their usefulness here; probably too expensive to train on the dataset and depending on class balance (not sure here what that is like) could have a bad balance of exploration/exploitation.</p>\n\n<p>Anyways to respond to your actual question: GPs are not a score but they can be used to find the minimum/maximum of an unknown function (such as the score of an ml model given hyperparameters) very quickly.  It may be tempting to assume that the GP has identified the minimum/maximum score for the ml model and what hyperparameters give that score, but you need to fall back on your CV scores to ultimately determine what hyperparameters you are using.  My personal experience with this is that I chose my model hyperparameters for final submission based on what the GP had as earning me the maximum score; however, it would have been much better to the model that had the best local CV (I dropped like 42% of the entire leaderboard because of this mistake).</p>\n\n<p>Sorry for the ambiguity.  You are totally right that CV scores and GPs can't be directly compared.</p>\n\n<p>You're right: I'm new here.  Looking forward to learning more from people who are way smarter than me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434263,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-06T05:55:32.687000",
          "content": "<p>OK, you mean using GP to optimize hyperparameters.  This is a very legitimate use case.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 432731,
      "author_name": "dylonLL",
      "author_url": "",
      "post_date": "2018-12-04T09:02:17.643000",
      "content": "<p>I have tried this one but it takes too long to fit just the training set so i dropped the idea</p>",
      "votes": 1,
      "replies": [
        {
          "id": 432753,
          "author_name": "Mithrillion",
          "author_url": "",
          "post_date": "2018-12-04T09:29:26.543000",
          "content": "<p>Yeah, GP is pretty slow (typically n^3). I think the other flaw is that it does not model the unseen portion of the data very well, so you can expect all sorts of weirdness being fitted if you have a supernova with only the rising phase. It is used with success in many previous researches, but from what I have read, they all had much more complete light curves to work with. I think a good curve model should at least be moderately resistant to partially obscuring the light curve, whereas samples drawn from a GP are usually completely dependent on the observed points for a given object, so it might not fare well with the particualr dataset we have.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 432885,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-04T12:36:34.363000",
          "content": "<blockquote>\n  <p>it does not model the unseen portion of the data very well</p>\n</blockquote>\n\n<p>Do you think autoencoders model unseen part better?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432901,
          "author_name": "Mithrillion",
          "author_url": "",
          "post_date": "2018-12-04T12:53:42.557000",
          "content": "<p>Not necessarily so. It'd be much easier if we had \"complete\" light curves as reference, in which case it becomes a curve restoration task. Piecing together broken light curves is much harder, and I still don't have a great solution for that. I can get an autoencoder to default to an \"average\" light curve rather than the mean with insufficient data, but extrapolation is still questionable. At least I was able to mostly eliminate instant drop to zero or straight line slope down to the next value now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435907,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2018-12-09T03:31:09.060000",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>Seriously, how do you manage to learn all this so fast? You wrote <code>having no clue about it</code> 5 days ago and now finishing test run ... I feel like damn blonde now...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 435950,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-09T06:11:23.580000",
          "content": "<p>This is funny, I always think I am a slow learner.  There is so much I don't know, and so much new things produced constantly, especially in deep learning.</p>\n\n<p>To compensate I look for relevant material and read it carefully.  For instance, how many of us here noticed the comment about celerite being faster than other gp packages in what we got from organizers?</p>\n\n<p>Back to gp, sure, I can compute stuff, but making good use of it is another story.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 433911,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-05T16:37:27.927000",
      "content": "<p>Looks pretty to me.  But it is not that pretty in many cases, still experimenting.  And being pretty does not mean useful either.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/433911/10832/pg.png\" alt=\"gaussian process\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 434021,
          "author_name": "Eavan Pattie",
          "author_url": "",
          "post_date": "2018-12-05T19:34:40.207000",
          "content": "<p>Thanks for this. How much time does it take to for you to compute this for the training set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434074,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-12-05T21:11:29.457000",
          "content": "<p>Hey CPMP - alot of the measurements look like they were done over 3 time periods - just saying if you can partition the data you will really speed up your Gaussian Processing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434125,
          "author_name": "Mithrillion",
          "author_url": "",
          "post_date": "2018-12-05T23:31:02.817000",
          "content": "<p>But then you don't get the benefit of interpolating unseen values...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434236,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-06T04:50:41.173000",
          "content": "<blockquote>\n  <p>How much time does it take to for you to compute this for the training set?</p>\n</blockquote>\n\n<p>About 1h30 with a single thread.  it means it can be computed in an affordable time on test data with some parallelism.</p>\n\n<blockquote>\n  <p>if you can partition the data you will really speed up your Gaussian Processing</p>\n</blockquote>\n\n<p>As @Mithrillion wrote, this defeats the purpose.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434632,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-06T17:48:12.567000",
          "content": "<p>This example looks like very good data example, not all objects are like that. </p>\n\n<p>BTW About 1h30 with a single thread -- it's on CPU, correct? Out of curiosity, how long do you think it will run for the test ? I've already paralleled periods on 24 CPUs and had to drop the idea after few days of running... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434641,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-06T18:06:50.797000",
          "content": "<p>Using 20 threads it should take a coupe of days.  I am still experimenting on how to best use the posterior of the gp.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435754,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-08T17:34:20.780000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> 3 days passed, did you try it on test, did it help your model ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435758,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-08T17:58:19.060000",
          "content": "<p>I am computing on test data, will see result tomorrow I think.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435905,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-09T03:27:33.810000",
          "content": "<p>Ok, I'll write down your current LB :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436220,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-09T23:38:50.107000",
          "content": "<p>LB did not change... so probably making use of GP is still problematic... is it ? We still need to find out how to get another 0.1 gain, time is shrinking...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436638,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-10T17:09:47.423000",
          "content": "<p>We're busy trying stuff.  Ensembling can wait a bit...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436659,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-10T17:56:23.627000",
          "content": "<p>So did it help on single model ? (I know you don't have to answer) :)  We are busy running instances, so trying to plan computational resources.... next competition I choose with the small test size! <br>\n(My God, I got addicted already, when it's not too late to quit :-)))?) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 436668,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-12-10T18:05:40.207000",
          "content": "<p>Wait until the Kaggle Blues kick in once the competition finishes! ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436670,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-10T18:12:20.840000",
          "content": "<p>@Blonde, I shared a lot already.  I am fine sharing pointer to publicly available information, but I won't tell you what works and what does not work.  Why don't you ask the leaders?  ;)</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 436672,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-10T18:14:38.033000",
          "content": "<p><a href=\"/scirpus\">@scirpus</a> </p>\n\n<p>I am actually on maternity, my Kaggle Blues is treating my Babies Blues... very effective treatment I must say :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 435542,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T07:30:56.700000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 435661,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-08T13:44:30.873000",
          "content": "<p>I thought I recommended a textbook in my post...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "432624": "If you are like me, i.e. willing to give a try at gaussian process modeling, but having no clue about it, then you may want to read this series of introductory posts: https://allofyourbases.com/category/gaussian-process-modelling/\n\nThe bible is this book, available online: http://www.gaussianprocess.org/gpml/chapters/\n\nThere is also this paper: http://www.robots.ox.ac.uk/~sjrob/Pubs/philTransA_2012.pdf\n\nGood read!",
    "438397": "Some nice gp fit\n![gp fit][1]\n\nThis also shows the limit of gaussian process modeling.  The peak was clearly before the second observation period, hence the extrapolation is a bit shaky.  But the curves look smoother that the autoencoder ones that have been shared so far.\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/438397/10873/gp.png",
    "435135": "From my own tests, GP is slow because you have inside a matrix inversion. Then you run it multiple time when optmizing the hyperparameters.\nI tested instead Kernel methods for time interpolation, and suprisingly achieved better Likelihood with the same number of parameters. (spread of the data around 0 when no input is present and time window length). Fitting only for the time length and setting the sigma to the std of the time serie  reachs almost as good likelihood, while running even faster.\n\nI agree with Mithrillon that having a template to extrapolate / interpolate lightcurve would help. However the exact template depends on the class... This is the research I'm doing now. For an easy class, you can look at my kernel for class 6:\nhttps://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6/comments ",
    "433020": "Is there anyone find GP useful?",
    "432731": "I have tried this one but it takes too long to fit just the training set so i dropped the idea",
    "435907": "@cpmpml\n\nSeriously, how do you manage to learn all this so fast? You wrote ```having no clue about it``` 5 days ago and now finishing test run ... I feel like damn blonde now...",
    "433911": "Looks pretty to me.  But it is not that pretty in many cases, still experimenting.  And being pretty does not mean useful either.\n\n![gaussian process][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/433911/10832/pg.png",
    "435542": ""
  }
}