{
  "id": 6694,
  "title": "Weather files from GRIB files",
  "url": "/competitions/flight2-main/discussion/6694",
  "author_name": "",
  "post_date": "2013-12-27T12:28:35.817Z",
  "votes": null,
  "comment_count": 18,
  "views": 8419,
  "content": "<p>Is there a straightforward way to extract the simulator weather files from GRIB files? In the milestone forum I found the link to the F# Grib Decoder (https://github.com/benhamner/GribDotNet/tree/master/GribDotNet), but it's functionality seems to be limited to create JPEGs from GRIB files. It's not clear to me how to extract the wind data from the byte-blocks of the raw data.</p>",
  "messages": [
    {
      "id": "36695",
      "postDate": "12/27/2013 12:28:35",
      "content": "<p>Is there a straightforward way to extract the simulator weather files from GRIB files? In the milestone forum I found the link to the F# Grib Decoder (https://github.com/benhamner/GribDotNet/tree/master/GribDotNet), but it's functionality seems to be limited to create JPEGs from GRIB files. It's not clear to me how to extract the wind data from the byte-blocks of the raw data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36724",
      "postDate": "12/27/2013 19:04:57",
      "content": "<p>No straightforward way that I know of. But you can get clues from Summary.fs.</p>\n<p>Here's my workflow:</p>\n<p>- Open GRIB file with GribReader.readGribFromPath</p>\n<p>- Filter the records to include just the ones containing UComponentOfWind and VComponentOfWind</p>\n<p>- Filter based on ProductDefinitionSection.ProductDefinitionTemplate (must be Type0)</p>\n<p>- Filter based on TypeOfFirstFixedSurface=HybridLevel</p>\n<p>- Filter based on altitude (select 15*2 items starting from index 36)</p>\n<p>- Use JpegDecoder.bitmapToGrid to decode the data</p>\n<p>- Transpose to obtain the 451x337 matrix you need</p>\n<p>- Convert the values from m/s to Knots</p>\n<p>- Save to weather.txt file</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36731",
      "postDate": "12/27/2013 20:32:06",
      "content": "<p>Missed a step. You also have to interleave the U and V components before writing to weather.txt.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36737",
      "postDate": "12/27/2013 22:06:35",
      "content": "<p>@Anil: Thank you very much for the helpful advices. I hope that I have now some correct code to convert the data.</p>\n<p>&nbsp;</p>\n<p>For all others of you working on the same issue, maybe one more useful hint: Check the file JpegDecoderTests.fs. It contains more or less all the relevant code to implement the filters mentioned above and to decode the image to a set of matrices. You just need to transpose them, convert from m/s to knots, and create a correct output format!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36738",
      "postDate": "12/27/2013 23:25:44",
      "content": "<p>[quote=Anil Thomas;36724]</p>\n<p>No straightforward way that I know of. But you can get clues from Summary.fs.</p>\n<p>Here's my workflow:</p>\n<p>- Open GRIB file with GribReader.readGribFromPath</p>\n<p>- Filter the records to include just the ones containing UComponentOfWind and VComponentOfWind</p>\n<p>- Filter based on ProductDefinitionSection.ProductDefinitionTemplate (must be Type0)</p>\n<p>- Filter based on TypeOfFirstFixedSurface=HybridLevel</p>\n<p>- Filter based on altitude (select 15*2 items starting from index 36)</p>\n<p>- Use JpegDecoder.bitmapToGrid to decode the data</p>\n<p>- Transpose to obtain the 451x337 matrix you need</p>\n<p>- Convert the values from m/s to Knots</p>\n<p>- Save to weather.txt file</p>\n<p>[/quote]</p>\n<p>Thanks for sharing Anil!</p>\n<p>Can you exactly recreate the weather file as the oneday data? I'm using pygrib (as I'm using only Python) and I've noticed that I have variations between +0.5 and -0.5 on all the wind components. Sounds like some rounding done at done at some points, but I haven't figured out where this is taking place!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36739",
      "postDate": "12/28/2013 00:52:29",
      "content": "<p>[quote=Alessandro Mariani;36738]</p>\n<p>Can you exactly recreate the weather file as the oneday data?</p>\n<p>[/quote]</p>\n<p>Not perfectly, but very close. From cursory eyeballing, I see an occasional deviation of +/- 0.01.</p>\n<p>[quote]</p>\n<p>I'm using pygrib (as I'm using only Python) and I've noticed that I have variations between +0.5 and -0.5 on all the wind components. </p>\n<p>[/quote]</p>\n<p>Maybe some mix-up in the grib files, like using a forecast instead of the real weather? Or not using the data from the right altitude? Note that there is data for 50 different altitudes in the grib files, out of which 15 have to be selected.</p>\n<p>I had started off with Python too. I couldn't figure out an efficient way to communicate with the simulator, so finally bit the bullet and learned to program in F#. Haven't developed any love for the language yet.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36740",
      "postDate": "12/28/2013 01:12:12",
      "content": "<p>Well... +/- 0.01 variation it's pretty much exact!</p>\n<p>While perturbation of +/- 0.5 are quite annoying sometimes unfortunately, and I'm using the snapshot file on the correct layers.. the difference is too homogeneous to be something else than rounding. I started off in F#, giving up after a while as I couldn't understand quite few things in the GribDotNet library, ending up installing cygwin to get pygrib running!</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36750",
      "postDate": "12/28/2013 11:36:46",
      "content": "<p>The rounding is happening in JPEG2000 extraction layer. The .NET code does extract a &quot;lightness&quot; component from the image, which gets returned in 8 bit precision, whereas the original data could be in 10 bits precision, for example.</p>\n<p>So, if you will use exactly the same code to extract values, you will get +/- 0.01 values. However, if you used a more precise implementation of the decoder (without rounding lightness to 8 bit), you could get errors as much as 0.5 or more.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36754",
      "postDate": "12/28/2013 14:22:26",
      "content": "<p>hmm... following your hint, I've converted the grib record (message\\layer however is called) from the original 10bits to 8bits. I'm getting different numbers, but still quite far away from the original :(</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36759",
      "postDate": "12/28/2013 15:03:38",
      "content": "<p>I did some testing a while ago and at the time it looked like the projection grid changed in some way during the JPEG2000 extraction (like skewing/translating) but I didn't pursue the issue very well, so it may be some kind of mistake on my part...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36760",
      "postDate": "12/28/2013 15:09:38",
      "content": "<p>It sound like I have to get my hands dirty again with F#! Pretty annoying as would be nice &amp; fair to have a script to generate these files without spending too much time on it...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37091",
      "postDate": "01/07/2014 20:15:02",
      "content": "<p>Does anyone know which grib files kaggle uses to extract the 8 snapshots for each test day on public leaderboard?&nbsp;</p>\n<p>For example for day 130911_1824</p>\n<p>does kaggle use&nbsp;</p>\n<p>13091118.rap.t18z.awp130bgrbf00.grib2<br>13091118.rap.t18z.awp130bgrbf01.grib2<br>13091118.rap.t18z.awp130bgrbf02.grib2<br>13091118.rap.t18z.awp130bgrbf03.grib2<br>13091118.rap.t18z.awp130bgrbf04.grib2<br>13091118.rap.t18z.awp130bgrbf05.grib2<br>13091118.rap.t18z.awp130bgrbf06.grib2<br>13091118.rap.t18z.awp130bgrbf07.grib2<br>13091118.rap.t18z.awp130bgrbf08.grib2</p>\n<p>or&nbsp;</p>\n<p>13091118.rap.t18z.awp130bgrbf00.grib2</p>\n<p>13091119.rap.t19z.awp130bgrbf00.grib2</p>\n<p>13091120.rap.t20z.awp130bgrbf00.grib2<br>13091121.rap.t21z.awp130bgrbf00.grib2<br>13091122.rap.t22z.awp130bgrbf00.grib2<br>13091123.rap.t23z.awp130bgrbf00.grib2<br>13091200.rap.t00z.awp130bgrbf00.grib2<br>13091201.rap.t01z.awp130bgrbf00.grib2</p>\n<p>thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37097",
      "postDate": "01/07/2014 22:27:26",
      "content": "<p>The second batch. But we're only allowed to use the first batch.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37106",
      "postDate": "01/08/2014 01:01:09",
      "content": "<p>So there are people who have downloaded the second batch of files and could have constructed a training environment similar to publicleaderboard using those files. But late participants like me who only have the first batch of files &nbsp;provided on the data page can never construct such training environment. Is it right?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37112",
      "postDate": "01/08/2014 05:02:44",
      "content": "<p>I think you're right. However, using &quot;second batch&quot; will only yield you a higher position on the current leaderboard, and on the final submission (February data) your model will have to use &quot;first batch&quot; or it will be disqualified.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37132",
      "postDate": "01/08/2014 15:29:34",
      "content": "<p>I understand that we are not supposed to use &nbsp;&quot;second batch&quot; for final submission. But is there anyone who has &quot;second batch&quot; for leaderboard test days kind enough to upload somewhere for us to download? Thanks so much.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37138",
      "postDate": "01/08/2014 16:53:13",
      "content": "<p>Nothing here unfortunately... =(</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37157",
      "postDate": "01/09/2014 07:57:02",
      "content": "<p>Here are the links to the files with &quot;second batch&quot; files for the current leaderboard data. I'm not sure if you can actually use these to post the submissions to the leaderboard (I'm not doing that), please check the rules yourself. But you could surely use this data (as it was public until NOAA has removed these from their site) to compare with the your predicted weather.</p>\n<p>http://goo.gl/1k3d9H</p>\n<p>http://goo.gl/Puw2l0</p>\n<p>http://goo.gl/45QUh8</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37168",
      "postDate": "01/09/2014 15:52:20",
      "content": "<p>Thanks,&nbsp;charango. I am not going to use it for prediction just for checking.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 36724,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "12/27/2013 19:04:57",
      "content": "<p>No straightforward way that I know of. But you can get clues from Summary.fs.</p>\n<p>Here's my workflow:</p>\n<p>- Open GRIB file with GribReader.readGribFromPath</p>\n<p>- Filter the records to include just the ones containing UComponentOfWind and VComponentOfWind</p>\n<p>- Filter based on ProductDefinitionSection.ProductDefinitionTemplate (must be Type0)</p>\n<p>- Filter based on TypeOfFirstFixedSurface=HybridLevel</p>\n<p>- Filter based on altitude (select 15*2 items starting from index 36)</p>\n<p>- Use JpegDecoder.bitmapToGrid to decode the data</p>\n<p>- Transpose to obtain the 451x337 matrix you need</p>\n<p>- Convert the values from m/s to Knots</p>\n<p>- Save to weather.txt file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36731,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "12/27/2013 20:32:06",
      "content": "<p>Missed a step. You also have to interleave the U and V components before writing to weather.txt.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36737,
      "author_name": "derdasmachenmuss",
      "author_url": "",
      "post_date": "12/27/2013 22:06:35",
      "content": "<p>@Anil: Thank you very much for the helpful advices. I hope that I have now some correct code to convert the data.</p>\n<p>&nbsp;</p>\n<p>For all others of you working on the same issue, maybe one more useful hint: Check the file JpegDecoderTests.fs. It contains more or less all the relevant code to implement the filters mentioned above and to decode the image to a set of matrices. You just need to transpose them, convert from m/s to knots, and create a correct output format!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36738,
      "author_name": "alzmcr",
      "author_url": "",
      "post_date": "12/27/2013 23:25:44",
      "content": "<p>[quote=Anil Thomas;36724]</p>\n<p>No straightforward way that I know of. But you can get clues from Summary.fs.</p>\n<p>Here's my workflow:</p>\n<p>- Open GRIB file with GribReader.readGribFromPath</p>\n<p>- Filter the records to include just the ones containing UComponentOfWind and VComponentOfWind</p>\n<p>- Filter based on ProductDefinitionSection.ProductDefinitionTemplate (must be Type0)</p>\n<p>- Filter based on TypeOfFirstFixedSurface=HybridLevel</p>\n<p>- Filter based on altitude (select 15*2 items starting from index 36)</p>\n<p>- Use JpegDecoder.bitmapToGrid to decode the data</p>\n<p>- Transpose to obtain the 451x337 matrix you need</p>\n<p>- Convert the values from m/s to Knots</p>\n<p>- Save to weather.txt file</p>\n<p>[/quote]</p>\n<p>Thanks for sharing Anil!</p>\n<p>Can you exactly recreate the weather file as the oneday data? I'm using pygrib (as I'm using only Python) and I've noticed that I have variations between +0.5 and -0.5 on all the wind components. Sounds like some rounding done at done at some points, but I haven't figured out where this is taking place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36739,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "12/28/2013 00:52:29",
      "content": "<p>[quote=Alessandro Mariani;36738]</p>\n<p>Can you exactly recreate the weather file as the oneday data?</p>\n<p>[/quote]</p>\n<p>Not perfectly, but very close. From cursory eyeballing, I see an occasional deviation of +/- 0.01.</p>\n<p>[quote]</p>\n<p>I'm using pygrib (as I'm using only Python) and I've noticed that I have variations between +0.5 and -0.5 on all the wind components. </p>\n<p>[/quote]</p>\n<p>Maybe some mix-up in the grib files, like using a forecast instead of the real weather? Or not using the data from the right altitude? Note that there is data for 50 different altitudes in the grib files, out of which 15 have to be selected.</p>\n<p>I had started off with Python too. I couldn't figure out an efficient way to communicate with the simulator, so finally bit the bullet and learned to program in F#. Haven't developed any love for the language yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36740,
      "author_name": "alzmcr",
      "author_url": "",
      "post_date": "12/28/2013 01:12:12",
      "content": "<p>Well... +/- 0.01 variation it's pretty much exact!</p>\n<p>While perturbation of +/- 0.5 are quite annoying sometimes unfortunately, and I'm using the snapshot file on the correct layers.. the difference is too homogeneous to be something else than rounding. I started off in F#, giving up after a while as I couldn't understand quite few things in the GribDotNet library, ending up installing cygwin to get pygrib running!</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36750,
      "author_name": "sergeykozub",
      "author_url": "",
      "post_date": "12/28/2013 11:36:46",
      "content": "<p>The rounding is happening in JPEG2000 extraction layer. The .NET code does extract a &quot;lightness&quot; component from the image, which gets returned in 8 bit precision, whereas the original data could be in 10 bits precision, for example.</p>\n<p>So, if you will use exactly the same code to extract values, you will get +/- 0.01 values. However, if you used a more precise implementation of the decoder (without rounding lightness to 8 bit), you could get errors as much as 0.5 or more.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36754,
      "author_name": "alzmcr",
      "author_url": "",
      "post_date": "12/28/2013 14:22:26",
      "content": "<p>hmm... following your hint, I've converted the grib record (message\\layer however is called) from the original 10bits to 8bits. I'm getting different numbers, but still quite far away from the original :(</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36759,
      "author_name": "gabrieleseppi",
      "author_url": "",
      "post_date": "12/28/2013 15:03:38",
      "content": "<p>I did some testing a while ago and at the time it looked like the projection grid changed in some way during the JPEG2000 extraction (like skewing/translating) but I didn't pursue the issue very well, so it may be some kind of mistake on my part...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36760,
      "author_name": "alzmcr",
      "author_url": "",
      "post_date": "12/28/2013 15:09:38",
      "content": "<p>It sound like I have to get my hands dirty again with F#! Pretty annoying as would be nice &amp; fair to have a script to generate these files without spending too much time on it...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37091,
      "author_name": "mlsword",
      "author_url": "",
      "post_date": "01/07/2014 20:15:02",
      "content": "<p>Does anyone know which grib files kaggle uses to extract the 8 snapshots for each test day on public leaderboard?&nbsp;</p>\n<p>For example for day 130911_1824</p>\n<p>does kaggle use&nbsp;</p>\n<p>13091118.rap.t18z.awp130bgrbf00.grib2<br>13091118.rap.t18z.awp130bgrbf01.grib2<br>13091118.rap.t18z.awp130bgrbf02.grib2<br>13091118.rap.t18z.awp130bgrbf03.grib2<br>13091118.rap.t18z.awp130bgrbf04.grib2<br>13091118.rap.t18z.awp130bgrbf05.grib2<br>13091118.rap.t18z.awp130bgrbf06.grib2<br>13091118.rap.t18z.awp130bgrbf07.grib2<br>13091118.rap.t18z.awp130bgrbf08.grib2</p>\n<p>or&nbsp;</p>\n<p>13091118.rap.t18z.awp130bgrbf00.grib2</p>\n<p>13091119.rap.t19z.awp130bgrbf00.grib2</p>\n<p>13091120.rap.t20z.awp130bgrbf00.grib2<br>13091121.rap.t21z.awp130bgrbf00.grib2<br>13091122.rap.t22z.awp130bgrbf00.grib2<br>13091123.rap.t23z.awp130bgrbf00.grib2<br>13091200.rap.t00z.awp130bgrbf00.grib2<br>13091201.rap.t01z.awp130bgrbf00.grib2</p>\n<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37097,
      "author_name": "gabrieleseppi",
      "author_url": "",
      "post_date": "01/07/2014 22:27:26",
      "content": "<p>The second batch. But we're only allowed to use the first batch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37106,
      "author_name": "mlsword",
      "author_url": "",
      "post_date": "01/08/2014 01:01:09",
      "content": "<p>So there are people who have downloaded the second batch of files and could have constructed a training environment similar to publicleaderboard using those files. But late participants like me who only have the first batch of files &nbsp;provided on the data page can never construct such training environment. Is it right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37112,
      "author_name": "sergeykozub",
      "author_url": "",
      "post_date": "01/08/2014 05:02:44",
      "content": "<p>I think you're right. However, using &quot;second batch&quot; will only yield you a higher position on the current leaderboard, and on the final submission (February data) your model will have to use &quot;first batch&quot; or it will be disqualified.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37132,
      "author_name": "mlsword",
      "author_url": "",
      "post_date": "01/08/2014 15:29:34",
      "content": "<p>I understand that we are not supposed to use &nbsp;&quot;second batch&quot; for final submission. But is there anyone who has &quot;second batch&quot; for leaderboard test days kind enough to upload somewhere for us to download? Thanks so much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37138,
      "author_name": "gabrieleseppi",
      "author_url": "",
      "post_date": "01/08/2014 16:53:13",
      "content": "<p>Nothing here unfortunately... =(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37157,
      "author_name": "sergeykozub",
      "author_url": "",
      "post_date": "01/09/2014 07:57:02",
      "content": "<p>Here are the links to the files with &quot;second batch&quot; files for the current leaderboard data. I'm not sure if you can actually use these to post the submissions to the leaderboard (I'm not doing that), please check the rules yourself. But you could surely use this data (as it was public until NOAA has removed these from their site) to compare with the your predicted weather.</p>\n<p>http://goo.gl/1k3d9H</p>\n<p>http://goo.gl/Puw2l0</p>\n<p>http://goo.gl/45QUh8</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37168,
      "author_name": "mlsword",
      "author_url": "",
      "post_date": "01/09/2014 15:52:20",
      "content": "<p>Thanks,&nbsp;charango. I am not going to use it for prediction just for checking.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "36695": "",
    "36724": "",
    "36731": "",
    "36737": "",
    "36738": "",
    "36739": "",
    "36740": "",
    "36750": "",
    "36754": "",
    "36759": "",
    "36760": "",
    "37091": "",
    "37097": "",
    "37106": "",
    "37112": "",
    "37132": "",
    "37138": "",
    "37157": "",
    "37168": ""
  },
  "source": "meta"
}