{
  "id": 71827,
  "title": "27 Test Set Basic Stats in 8 minutes",
  "url": "/competitions/PLAsTiCC-2018/discussion/71827",
  "author_name": "Mithrillion",
  "post_date": "2018-11-17T05:38:15.609000",
  "votes": 19,
  "comment_count": 9,
  "views": 0,
  "content": "<p>It's time to put skills learnt from a certain botched competition to use! The following BigQuery script extracts the same features as the kernel <a href=\"https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135\">https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135</a> minus the tsfresh features, but in just 8 minutes without having to manually split time series by objects.</p>\n\n<p><code>\nWITH <br>\n  aux AS(\n  SELECT\n    object_id,\n    POW(flux/flux_err, 2) AS flux_ratio_sq,\n    flux * POW(flux/flux_err, 2) AS flux_by_flux_ratio_sq\n  FROM\n    `project.astro.series` ), <br>\n  simple AS(\n  SELECT\n    src.object_id,\n    AVG(flux) AS flux_mean,\n    MAX(flux) AS flux_max,\n    MIN(flux) AS flux_min,\n    APPROX_QUANTILES(flux, 2)[ORDINAL(1)] AS flux_median,\n    STDDEV(flux) AS flux_std,\n    AVG(flux_err) AS flux_err_mean,\n    MAX(flux_err) AS flux_err_max,\n    MIN(flux_err) AS flux_err_min,\n    APPROX_QUANTILES(flux_err, 2)[ORDINAL(1)] AS flux_err_median,\n    STDDEV(flux_err) AS flux_err_std,\n    AVG(detected) AS detected_mean,\n    AVG(flux_ratio_sq) AS flux_ratio_sq_mean,\n    STDDEV(flux_ratio_sq) AS flux_ratio_sq_std,\n    SUM(flux_ratio_sq) AS flux_ratio_sq_sum,\n    AVG(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_mean,\n    STDDEV(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_std,\n    SUM(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_sum\n  FROM\n    `project.astro.series` src,\n    aux\n  WHERE\n    src.object_id = aux.object_id\n  GROUP BY\n    object_id ), <br>\n  skews AS (\n  SELECT\n    src.object_id,\n    AVG(POW((flux - flux_mean) / flux_std, 3)) AS flux_skew,\n    AVG(POW((flux_err - flux_mean) / flux_std, 3)) AS flux_err_skew,\n    AVG(POW((flux_ratio_sq - flux_ratio_sq_mean) / flux_ratio_sq_std, 3)) AS flux_ratio_sq_skew,\n    AVG(POW((flux_by_flux_ratio_sq - flux_by_flux_ratio_sq_mean) / flux_by_flux_ratio_sq_std, 3)) AS flux_by_flux_ratio_sq_skew\n  FROM\n    aux,\n    simple,\n    `project.astro.series` src\n  WHERE\n    aux.object_id = simple.object_id\n    AND aux.object_id = src.object_id\n  GROUP BY\n    object_id), <br>\n  det_mjd AS (\n  SELECT\n    object_id,\n    MAX(mjd) - MIN(mjd) AS det_mjd_diff\n  FROM\n    `project.astro.series` src\n  WHERE\n    detected = 1\n  GROUP BY\n    object_id ) <br>\nSELECT\n  simple.*,\n  flux_max - flux_min AS flux_diff,\n  (flux_max - flux_min) / flux_mean AS flux_diff2,\n  flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum AS flux_w_mean,\n  (flux_max - flux_min) / (flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum) AS flux_dif3,\n  flux_skew,\n  flux_err_skew,\n  flux_ratio_sq_skew,\n  flux_by_flux_ratio_sq_skew,\n  det_mjd_diff <br>\nFROM\n  simple, skews, det_mjd <br>\nWHERE\n  simple.object_id = skews.object_id\n  AND simple.object_id = det_mjd.object_id\n</code></p>\n\n<p><code>\nQuery complete (7 min 59.079 sec elapsed, 16.9 GB processed)\n</code></p>",
  "messages": [
    {
      "id": 422939,
      "postDate": "2018-11-17T05:38:15.610Z",
      "content": "<p>It's time to put skills learnt from a certain botched competition to use! The following BigQuery script extracts the same features as the kernel <a href=\"https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135\">https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135</a> minus the tsfresh features, but in just 8 minutes without having to manually split time series by objects.</p>\n\n<p><code>\nWITH <br>\n  aux AS(\n  SELECT\n    object_id,\n    POW(flux/flux_err, 2) AS flux_ratio_sq,\n    flux * POW(flux/flux_err, 2) AS flux_by_flux_ratio_sq\n  FROM\n    `project.astro.series` ), <br>\n  simple AS(\n  SELECT\n    src.object_id,\n    AVG(flux) AS flux_mean,\n    MAX(flux) AS flux_max,\n    MIN(flux) AS flux_min,\n    APPROX_QUANTILES(flux, 2)[ORDINAL(1)] AS flux_median,\n    STDDEV(flux) AS flux_std,\n    AVG(flux_err) AS flux_err_mean,\n    MAX(flux_err) AS flux_err_max,\n    MIN(flux_err) AS flux_err_min,\n    APPROX_QUANTILES(flux_err, 2)[ORDINAL(1)] AS flux_err_median,\n    STDDEV(flux_err) AS flux_err_std,\n    AVG(detected) AS detected_mean,\n    AVG(flux_ratio_sq) AS flux_ratio_sq_mean,\n    STDDEV(flux_ratio_sq) AS flux_ratio_sq_std,\n    SUM(flux_ratio_sq) AS flux_ratio_sq_sum,\n    AVG(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_mean,\n    STDDEV(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_std,\n    SUM(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_sum\n  FROM\n    `project.astro.series` src,\n    aux\n  WHERE\n    src.object_id = aux.object_id\n  GROUP BY\n    object_id ), <br>\n  skews AS (\n  SELECT\n    src.object_id,\n    AVG(POW((flux - flux_mean) / flux_std, 3)) AS flux_skew,\n    AVG(POW((flux_err - flux_mean) / flux_std, 3)) AS flux_err_skew,\n    AVG(POW((flux_ratio_sq - flux_ratio_sq_mean) / flux_ratio_sq_std, 3)) AS flux_ratio_sq_skew,\n    AVG(POW((flux_by_flux_ratio_sq - flux_by_flux_ratio_sq_mean) / flux_by_flux_ratio_sq_std, 3)) AS flux_by_flux_ratio_sq_skew\n  FROM\n    aux,\n    simple,\n    `project.astro.series` src\n  WHERE\n    aux.object_id = simple.object_id\n    AND aux.object_id = src.object_id\n  GROUP BY\n    object_id), <br>\n  det_mjd AS (\n  SELECT\n    object_id,\n    MAX(mjd) - MIN(mjd) AS det_mjd_diff\n  FROM\n    `project.astro.series` src\n  WHERE\n    detected = 1\n  GROUP BY\n    object_id ) <br>\nSELECT\n  simple.*,\n  flux_max - flux_min AS flux_diff,\n  (flux_max - flux_min) / flux_mean AS flux_diff2,\n  flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum AS flux_w_mean,\n  (flux_max - flux_min) / (flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum) AS flux_dif3,\n  flux_skew,\n  flux_err_skew,\n  flux_ratio_sq_skew,\n  flux_by_flux_ratio_sq_skew,\n  det_mjd_diff <br>\nFROM\n  simple, skews, det_mjd <br>\nWHERE\n  simple.object_id = skews.object_id\n  AND simple.object_id = det_mjd.object_id\n</code></p>\n\n<p><code>\nQuery complete (7 min 59.079 sec elapsed, 16.9 GB processed)\n</code></p>",
      "rawMarkdown": "It's time to put skills learnt from a certain botched competition to use! The following BigQuery script extracts the same features as the kernel https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135 minus the tsfresh features, but in just 8 minutes without having to manually split time series by objects.\n\n```\nWITH  \n  aux AS(\n  SELECT\n    object_id,\n    POW(flux/flux_err, 2) AS flux_ratio_sq,\n    flux * POW(flux/flux_err, 2) AS flux_by_flux_ratio_sq\n  FROM\n    `project.astro.series` ),  \n  simple AS(\n  SELECT\n    src.object_id,\n    AVG(flux) AS flux_mean,\n    MAX(flux) AS flux_max,\n    MIN(flux) AS flux_min,\n    APPROX_QUANTILES(flux, 2)[ORDINAL(1)] AS flux_median,\n    STDDEV(flux) AS flux_std,\n    AVG(flux_err) AS flux_err_mean,\n    MAX(flux_err) AS flux_err_max,\n    MIN(flux_err) AS flux_err_min,\n    APPROX_QUANTILES(flux_err, 2)[ORDINAL(1)] AS flux_err_median,\n    STDDEV(flux_err) AS flux_err_std,\n    AVG(detected) AS detected_mean,\n    AVG(flux_ratio_sq) AS flux_ratio_sq_mean,\n    STDDEV(flux_ratio_sq) AS flux_ratio_sq_std,\n    SUM(flux_ratio_sq) AS flux_ratio_sq_sum,\n    AVG(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_mean,\n    STDDEV(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_std,\n    SUM(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_sum\n  FROM\n    `project.astro.series` src,\n    aux\n  WHERE\n    src.object_id = aux.object_id\n  GROUP BY\n    object_id ),  \n  skews AS (\n  SELECT\n    src.object_id,\n    AVG(POW((flux - flux_mean) / flux_std, 3)) AS flux_skew,\n    AVG(POW((flux_err - flux_mean) / flux_std, 3)) AS flux_err_skew,\n    AVG(POW((flux_ratio_sq - flux_ratio_sq_mean) / flux_ratio_sq_std, 3)) AS flux_ratio_sq_skew,\n    AVG(POW((flux_by_flux_ratio_sq - flux_by_flux_ratio_sq_mean) / flux_by_flux_ratio_sq_std, 3)) AS flux_by_flux_ratio_sq_skew\n  FROM\n    aux,\n    simple,\n    `project.astro.series` src\n  WHERE\n    aux.object_id = simple.object_id\n    AND aux.object_id = src.object_id\n  GROUP BY\n    object_id),  \n  det_mjd AS (\n  SELECT\n    object_id,\n    MAX(mjd) - MIN(mjd) AS det_mjd_diff\n  FROM\n    `project.astro.series` src\n  WHERE\n    detected = 1\n  GROUP BY\n    object_id )  \nSELECT\n  simple.*,\n  flux_max - flux_min AS flux_diff,\n  (flux_max - flux_min) / flux_mean AS flux_diff2,\n  flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum AS flux_w_mean,\n  (flux_max - flux_min) / (flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum) AS flux_dif3,\n  flux_skew,\n  flux_err_skew,\n  flux_ratio_sq_skew,\n  flux_by_flux_ratio_sq_skew,\n  det_mjd_diff  \nFROM\n  simple, skews, det_mjd  \nWHERE\n  simple.object_id = skews.object_id\n  AND simple.object_id = det_mjd.object_id\n```\n\n```\nQuery complete (7 min 59.079 sec elapsed, 16.9 GB processed)\n```",
      "votes": 19
    },
    {
      "id": 442798,
      "postDate": "2018-12-20T14:17:51.853Z",
      "content": "<p><a href=\"/mithrillion\">@mithrillion</a>. Sorry for brining back this topic. I have not used BigQuery before. Could you direct me to a resource or tell me how this is used. During the competition, when you posted this, I had already calculated the features and stored it. But, now I would like to know how it is used. Can I run it on my local machine, or on the Kaggle kernel? I couldn't find a proper way to run this.  </p>",
      "rawMarkdown": "@mithrillion. Sorry for brining back this topic. I have not used BigQuery before. Could you direct me to a resource or tell me how this is used. During the competition, when you posted this, I had already calculated the features and stored it. But, now I would like to know how it is used. Can I run it on my local machine, or on the Kaggle kernel? I couldn't find a proper way to run this.  ",
      "replies": [
        {
          "id": 442802,
          "postDate": "2018-12-20T14:35:54.270Z",
          "content": "<p>Hi Vig Nam, no need to be sorry. Always glad to see new life in old threads. I think a good place to start is past competitions that host data on BigQuery. <a href=\"https://www.kaggle.com/juliaelliott/ga-bigquery-starter-kernel\">This</a> is an example of how to access public BigQuery datasets from a Kaggle kernel. For more advanced uses you might have to get a Google Cloud account and upload your own data to your BigQuery projects. It runs on the cloud, but there exists various APIs to directly access the data from the client side.</p>",
          "rawMarkdown": "Hi Vig Nam, no need to be sorry. Always glad to see new life in old threads. I think a good place to start is past competitions that host data on BigQuery. [This](https://www.kaggle.com/juliaelliott/ga-bigquery-starter-kernel) is an example of how to access public BigQuery datasets from a Kaggle kernel. For more advanced uses you might have to get a Google Cloud account and upload your own data to your BigQuery projects. It runs on the cloud, but there exists various APIs to directly access the data from the client side.",
          "votes": 1
        },
        {
          "id": 442831,
          "postDate": "2018-12-20T15:27:39.933Z",
          "content": "<p>Thanks for that info. Will check.</p>",
          "rawMarkdown": "Thanks for that info. Will check."
        }
      ]
    },
    {
      "id": 423123,
      "postDate": "2018-11-17T14:50:29.320Z",
      "content": "<p>Very nice! How much time do you estimate you are saving?</p>",
      "rawMarkdown": "Very nice! How much time do you estimate you are saving?",
      "replies": [
        {
          "id": 423130,
          "postDate": "2018-11-17T15:09:24.220Z",
          "content": "<p>I estimated that the feature extraction part (minus the fft etc.) of iprapas's kernel would take around 30-40 minutes to run on my i7 desktop. Since much of Pandas is not parallelised, it should perform similarly on processors with more (or less) cores. Maybe it can be reduced to 10-20 minutes with clever multiprocessing. However BigQuery can also be used to calculate slightly more complicated features like autocorrelation fairly fast, which might take much longer with Pandas, so I think the time you'll have to spend uploading / downloading data should be worth it.</p>",
          "rawMarkdown": "I estimated that the feature extraction part (minus the fft etc.) of iprapas's kernel would take around 30-40 minutes to run on my i7 desktop. Since much of Pandas is not parallelised, it should perform similarly on processors with more (or less) cores. Maybe it can be reduced to 10-20 minutes with clever multiprocessing. However BigQuery can also be used to calculate slightly more complicated features like autocorrelation fairly fast, which might take much longer with Pandas, so I think the time you'll have to spend uploading / downloading data should be worth it."
        }
      ]
    },
    {
      "id": 423017,
      "postDate": "2018-11-17T09:18:41.287Z",
      "content": "<p>That's a very elegant way to do it, thanks for sharing</p>",
      "rawMarkdown": "That's a very elegant way to do it, thanks for sharing"
    },
    {
      "id": 425003,
      "postDate": "2018-11-21T01:48:41.880Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 423170,
      "postDate": "2018-11-17T17:04:14.037Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 423319,
      "postDate": "2018-11-17T22:58:58.333Z",
      "content": "<p>Nice job, thanks a lot !</p>",
      "rawMarkdown": "Nice job, thanks a lot !"
    }
  ],
  "comments": [
    {
      "id": 442798,
      "author_name": "Vig",
      "author_url": "",
      "post_date": "2018-12-20T14:17:51.853000",
      "content": "<p><a href=\"/mithrillion\">@mithrillion</a>. Sorry for brining back this topic. I have not used BigQuery before. Could you direct me to a resource or tell me how this is used. During the competition, when you posted this, I had already calculated the features and stored it. But, now I would like to know how it is used. Can I run it on my local machine, or on the Kaggle kernel? I couldn't find a proper way to run this.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 442802,
          "author_name": "Mithrillion",
          "author_url": "",
          "post_date": "2018-12-20T14:35:54.270000",
          "content": "<p>Hi Vig Nam, no need to be sorry. Always glad to see new life in old threads. I think a good place to start is past competitions that host data on BigQuery. <a href=\"https://www.kaggle.com/juliaelliott/ga-bigquery-starter-kernel\">This</a> is an example of how to access public BigQuery datasets from a Kaggle kernel. For more advanced uses you might have to get a Google Cloud account and upload your own data to your BigQuery projects. It runs on the cloud, but there exists various APIs to directly access the data from the client side.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 442831,
          "author_name": "Vig",
          "author_url": "",
          "post_date": "2018-12-20T15:27:39.933000",
          "content": "<p>Thanks for that info. Will check.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423123,
      "author_name": "S D",
      "author_url": "",
      "post_date": "2018-11-17T14:50:29.320000",
      "content": "<p>Very nice! How much time do you estimate you are saving?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 423130,
          "author_name": "Mithrillion",
          "author_url": "",
          "post_date": "2018-11-17T15:09:24.220000",
          "content": "<p>I estimated that the feature extraction part (minus the fft etc.) of iprapas's kernel would take around 30-40 minutes to run on my i7 desktop. Since much of Pandas is not parallelised, it should perform similarly on processors with more (or less) cores. Maybe it can be reduced to 10-20 minutes with clever multiprocessing. However BigQuery can also be used to calculate slightly more complicated features like autocorrelation fairly fast, which might take much longer with Pandas, so I think the time you'll have to spend uploading / downloading data should be worth it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423017,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "2018-11-17T09:18:41.287000",
      "content": "<p>That's a very elegant way to do it, thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425003,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-21T01:48:41.880000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423170,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-17T17:04:14.037000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423319,
      "author_name": "mezoganet",
      "author_url": "",
      "post_date": "2018-11-17T22:58:58.333000",
      "content": "<p>Nice job, thanks a lot !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422939": "It's time to put skills learnt from a certain botched competition to use! The following BigQuery script extracts the same features as the kernel https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135 minus the tsfresh features, but in just 8 minutes without having to manually split time series by objects.\n\n```\nWITH  \n  aux AS(\n  SELECT\n    object_id,\n    POW(flux/flux_err, 2) AS flux_ratio_sq,\n    flux * POW(flux/flux_err, 2) AS flux_by_flux_ratio_sq\n  FROM\n    `project.astro.series` ),  \n  simple AS(\n  SELECT\n    src.object_id,\n    AVG(flux) AS flux_mean,\n    MAX(flux) AS flux_max,\n    MIN(flux) AS flux_min,\n    APPROX_QUANTILES(flux, 2)[ORDINAL(1)] AS flux_median,\n    STDDEV(flux) AS flux_std,\n    AVG(flux_err) AS flux_err_mean,\n    MAX(flux_err) AS flux_err_max,\n    MIN(flux_err) AS flux_err_min,\n    APPROX_QUANTILES(flux_err, 2)[ORDINAL(1)] AS flux_err_median,\n    STDDEV(flux_err) AS flux_err_std,\n    AVG(detected) AS detected_mean,\n    AVG(flux_ratio_sq) AS flux_ratio_sq_mean,\n    STDDEV(flux_ratio_sq) AS flux_ratio_sq_std,\n    SUM(flux_ratio_sq) AS flux_ratio_sq_sum,\n    AVG(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_mean,\n    STDDEV(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_std,\n    SUM(flux_by_flux_ratio_sq) AS flux_by_flux_ratio_sq_sum\n  FROM\n    `project.astro.series` src,\n    aux\n  WHERE\n    src.object_id = aux.object_id\n  GROUP BY\n    object_id ),  \n  skews AS (\n  SELECT\n    src.object_id,\n    AVG(POW((flux - flux_mean) / flux_std, 3)) AS flux_skew,\n    AVG(POW((flux_err - flux_mean) / flux_std, 3)) AS flux_err_skew,\n    AVG(POW((flux_ratio_sq - flux_ratio_sq_mean) / flux_ratio_sq_std, 3)) AS flux_ratio_sq_skew,\n    AVG(POW((flux_by_flux_ratio_sq - flux_by_flux_ratio_sq_mean) / flux_by_flux_ratio_sq_std, 3)) AS flux_by_flux_ratio_sq_skew\n  FROM\n    aux,\n    simple,\n    `project.astro.series` src\n  WHERE\n    aux.object_id = simple.object_id\n    AND aux.object_id = src.object_id\n  GROUP BY\n    object_id),  \n  det_mjd AS (\n  SELECT\n    object_id,\n    MAX(mjd) - MIN(mjd) AS det_mjd_diff\n  FROM\n    `project.astro.series` src\n  WHERE\n    detected = 1\n  GROUP BY\n    object_id )  \nSELECT\n  simple.*,\n  flux_max - flux_min AS flux_diff,\n  (flux_max - flux_min) / flux_mean AS flux_diff2,\n  flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum AS flux_w_mean,\n  (flux_max - flux_min) / (flux_by_flux_ratio_sq_sum / flux_ratio_sq_sum) AS flux_dif3,\n  flux_skew,\n  flux_err_skew,\n  flux_ratio_sq_skew,\n  flux_by_flux_ratio_sq_skew,\n  det_mjd_diff  \nFROM\n  simple, skews, det_mjd  \nWHERE\n  simple.object_id = skews.object_id\n  AND simple.object_id = det_mjd.object_id\n```\n\n```\nQuery complete (7 min 59.079 sec elapsed, 16.9 GB processed)\n```",
    "442798": "@mithrillion. Sorry for brining back this topic. I have not used BigQuery before. Could you direct me to a resource or tell me how this is used. During the competition, when you posted this, I had already calculated the features and stored it. But, now I would like to know how it is used. Can I run it on my local machine, or on the Kaggle kernel? I couldn't find a proper way to run this.  ",
    "423123": "Very nice! How much time do you estimate you are saving?",
    "423017": "That's a very elegant way to do it, thanks for sharing",
    "425003": "",
    "423170": "",
    "423319": "Nice job, thanks a lot !"
  }
}