{
  "id": 27926,
  "title": "4th place solution",
  "url": "/competitions/outbrain-click-prediction/writeups/andrii-cherednychenko-4th-place-solution",
  "author_name": "",
  "post_date": "2017-01-19T16:19:37.383Z",
  "votes": 30,
  "comment_count": 6,
  "views": 319,
  "content": "<p>Hey there, </p>\n\n<p>Congrats to the winners and thanks a lot for congrats to my side :-)</p>\n\n<p>Sharing details on my solution if anyone is interested, basically it is FFM model averaged for last 5 runs:</p>\n\n<p>Technologies / frameworks:</p>\n\n<ol>\n<li><p>Early merging of files was done using Python. Basically, I merged all files into  train/test/sumission files, then was reading line by line for generating features</p></li>\n<li><p>All feature generation and statistics building is done using Scala. I found it way faster than Python + I am better with it, especially when it came to parallelization. \nI have also filtered page_views.csv to users seen in train/submission sets, so different statistics would fit in RAM. Size of FFM train file ~180G. RAM required ~ 80G</p></li>\n<li><p>I have not used hashing for features, each feature seen &gt;=10 times would get its own ID. Submission features unseen features in train are dropped</p></li>\n<li><p>For numerical features, I have different strategies of dealing with them - bins, log smoothing or leaving as is, best chosen for each, based on validation. Some fields contain single feature, some bag of features. Total number of features is &gt;5M</p></li>\n<li><p>I added MEAP metric into libffm, so it would stop based on it, not on logloss. Used default parameters of libffm, including k=4</p></li>\n<li><p>No special approach was taken to deal with leak, removing  it reduces model's precision + model learned almost all leaks naturally, so no post processing was required</p></li>\n<li><p>Most of the features engineering was done on 32Gb laptop, later I used 128Gb RAM  machine and for the last week used AWS</p></li>\n<li><p>What have not worked for me - GBT leaves as features, users clustering , MF of page_views.csv, stacking of FFM. Also, run out of time for stacking ffm-&gt;xgboost. </p></li>\n<li><p>List of features:</p></li>\n</ol>\n\n<blockquote>\n  <p>country, state, platform, county, pageDocumentCategories,\n  pageDocumentEntities, pageDocumentTopics, publisherCTRAdv,\n  publisherCTRCamp, publisherCTRDoc, countryCTRAdv, countryCTRCamp,\n  countryCTRDoc, stateCTRAdv, stateCTRCamp, stateCTRDoc, countyCTRAdv,\n  countyCTRCamp, countyCTRDoc, prevDayCTRAdv, prevDayCTRCamp,\n  prevDayCTRDoc, currentDayCTRAdv, currentDayCTRCamp, currentDayCTRDoc,\n  nextDayCTRAdv, nextDayCTRCamp, nextDayCTRDoc, currentHourCTRAdv,\n  currentHourCTRCamp, currentHourCTRDoc, prevHourCTRAdv,\n  prevHourCTRCamp, prevHourCTRDoc, nextHourCTRAdv, nextHourCTRCamp,\n  nextHourCTRDoc, publisherSourceCTRAdv, publisherSourceCTRCamp,\n  publisherSourceCTRDoc, adId, documentId, campaignId, advertiserId,\n  userPageViewsMeta1, userPageViewsMeta2, documentIdViews,\n  campaignIdViews, advertiserIdViews, userDocsSeenFromLogToday,\n  userPageMeta1SeenToday, userPageMeta2SeenToday,\n  userCampSeenFromLogToday, userAdvertisersSeenFromLogToday,\n  userDocsSeenFromLogTomorrow, userPageMeta1SeenTomorrow,\n  userPageMeta2SeenTomorrow, docStats, documentCategoriesId, advStats,\n  documentEntitiesId, documentTopicsId, metaSourceId, metaPublisherId,\n  pageDocumentMeta1, pageDocumentMeta2, userClickedThisDocumentTimes,\n  userSkippedThisDocumentTimes, userClickedThisCampaignTimes,\n  userSkippedThisCampaignTimes, userClickedThisAdvertiserTimes,\n  userSkippedThisAdvertiserTimes, hourOfDay, dayNum, dayOfWeek,\n  thisAdEntityClickedBefore, daysFromAdDocPublished,\n  thisAdEntityClickedBeforeCount, thisAdCategoryClickedBefore,\n  thisAdCategoryClickedBeforeCount, userClickedThisAdTimes,\n  userSkippedThisAdTimes, userPageViewsDocuments,\n  userPageViewsCampaigns, userPageViewsAdvertisers,\n  userPageViewsCategories, userPageViewsEntitites, userPageViewsTopics,\n  seenThisCategoryInLog, userSeenThisMeta1, userSeenThisMeta2,\n  seenThisDocInLog, seenThisCampaignInLog, seenThisAdvInLog, campStats,\n  adStats, userCampSeenFromLogTomorrow,\n  userAdvertisersSeenFromLogTomorrow, seenThisEntityInLog,\n  userAllAdsFreq, nextDocUserClicked, nextCampUserClicked,\n  nextAdvUserClicked, userClickedCount, userSkippedCount,\n  userClickedCountToday, userDocsClickedToday, userCampClickedToday,\n  userAdvertisersClickedToday, seenThisTopicInLog,\n  userClickedThisDocFreq, userClickedThisCampFreq,\n  userClickedThisAdvFreq, lastPageViewUserSeen, nextPageViewUserSeen,\n  lastDocUserClicked, lastCampUserClicked, lastAdvUserClicked</p>\n</blockquote>\n\n<p>Thanks,</p>\n\n<p>Andrii</p>",
  "messages": [
    {
      "id": "157193",
      "postDate": "01/19/2017 16:19:37",
      "content": "<p>Hey there, </p>\n\n<p>Congrats to the winners and thanks a lot for congrats to my side :-)</p>\n\n<p>Sharing details on my solution if anyone is interested, basically it is FFM model averaged for last 5 runs:</p>\n\n<p>Technologies / frameworks:</p>\n\n<ol>\n<li><p>Early merging of files was done using Python. Basically, I merged all files into  train/test/sumission files, then was reading line by line for generating features</p></li>\n<li><p>All feature generation and statistics building is done using Scala. I found it way faster than Python + I am better with it, especially when it came to parallelization. \nI have also filtered page_views.csv to users seen in train/submission sets, so different statistics would fit in RAM. Size of FFM train file ~180G. RAM required ~ 80G</p></li>\n<li><p>I have not used hashing for features, each feature seen &gt;=10 times would get its own ID. Submission features unseen features in train are dropped</p></li>\n<li><p>For numerical features, I have different strategies of dealing with them - bins, log smoothing or leaving as is, best chosen for each, based on validation. Some fields contain single feature, some bag of features. Total number of features is &gt;5M</p></li>\n<li><p>I added MEAP metric into libffm, so it would stop based on it, not on logloss. Used default parameters of libffm, including k=4</p></li>\n<li><p>No special approach was taken to deal with leak, removing  it reduces model's precision + model learned almost all leaks naturally, so no post processing was required</p></li>\n<li><p>Most of the features engineering was done on 32Gb laptop, later I used 128Gb RAM  machine and for the last week used AWS</p></li>\n<li><p>What have not worked for me - GBT leaves as features, users clustering , MF of page_views.csv, stacking of FFM. Also, run out of time for stacking ffm-&gt;xgboost. </p></li>\n<li><p>List of features:</p></li>\n</ol>\n\n<blockquote>\n  <p>country, state, platform, county, pageDocumentCategories,\n  pageDocumentEntities, pageDocumentTopics, publisherCTRAdv,\n  publisherCTRCamp, publisherCTRDoc, countryCTRAdv, countryCTRCamp,\n  countryCTRDoc, stateCTRAdv, stateCTRCamp, stateCTRDoc, countyCTRAdv,\n  countyCTRCamp, countyCTRDoc, prevDayCTRAdv, prevDayCTRCamp,\n  prevDayCTRDoc, currentDayCTRAdv, currentDayCTRCamp, currentDayCTRDoc,\n  nextDayCTRAdv, nextDayCTRCamp, nextDayCTRDoc, currentHourCTRAdv,\n  currentHourCTRCamp, currentHourCTRDoc, prevHourCTRAdv,\n  prevHourCTRCamp, prevHourCTRDoc, nextHourCTRAdv, nextHourCTRCamp,\n  nextHourCTRDoc, publisherSourceCTRAdv, publisherSourceCTRCamp,\n  publisherSourceCTRDoc, adId, documentId, campaignId, advertiserId,\n  userPageViewsMeta1, userPageViewsMeta2, documentIdViews,\n  campaignIdViews, advertiserIdViews, userDocsSeenFromLogToday,\n  userPageMeta1SeenToday, userPageMeta2SeenToday,\n  userCampSeenFromLogToday, userAdvertisersSeenFromLogToday,\n  userDocsSeenFromLogTomorrow, userPageMeta1SeenTomorrow,\n  userPageMeta2SeenTomorrow, docStats, documentCategoriesId, advStats,\n  documentEntitiesId, documentTopicsId, metaSourceId, metaPublisherId,\n  pageDocumentMeta1, pageDocumentMeta2, userClickedThisDocumentTimes,\n  userSkippedThisDocumentTimes, userClickedThisCampaignTimes,\n  userSkippedThisCampaignTimes, userClickedThisAdvertiserTimes,\n  userSkippedThisAdvertiserTimes, hourOfDay, dayNum, dayOfWeek,\n  thisAdEntityClickedBefore, daysFromAdDocPublished,\n  thisAdEntityClickedBeforeCount, thisAdCategoryClickedBefore,\n  thisAdCategoryClickedBeforeCount, userClickedThisAdTimes,\n  userSkippedThisAdTimes, userPageViewsDocuments,\n  userPageViewsCampaigns, userPageViewsAdvertisers,\n  userPageViewsCategories, userPageViewsEntitites, userPageViewsTopics,\n  seenThisCategoryInLog, userSeenThisMeta1, userSeenThisMeta2,\n  seenThisDocInLog, seenThisCampaignInLog, seenThisAdvInLog, campStats,\n  adStats, userCampSeenFromLogTomorrow,\n  userAdvertisersSeenFromLogTomorrow, seenThisEntityInLog,\n  userAllAdsFreq, nextDocUserClicked, nextCampUserClicked,\n  nextAdvUserClicked, userClickedCount, userSkippedCount,\n  userClickedCountToday, userDocsClickedToday, userCampClickedToday,\n  userAdvertisersClickedToday, seenThisTopicInLog,\n  userClickedThisDocFreq, userClickedThisCampFreq,\n  userClickedThisAdvFreq, lastPageViewUserSeen, nextPageViewUserSeen,\n  lastDocUserClicked, lastCampUserClicked, lastAdvUserClicked</p>\n</blockquote>\n\n<p>Thanks,</p>\n\n<p>Andrii</p>",
      "rawMarkdown": "Hey there, \r\n\r\nCongrats to the winners and thanks a lot for congrats to my side :-)\r\n\r\n\r\nSharing details on my solution if anyone is interested, basically it is FFM model averaged for last 5 runs:\r\n\r\nTechnologies / frameworks:\r\n\r\n1. Early merging of files was done using Python. Basically, I merged all files into  train/test/sumission files, then was reading line by line for generating features\r\n\r\n2. All feature generation and statistics building is done using Scala. I found it way faster than Python + I am better with it, especially when it came to parallelization. \r\nI have also filtered page_views.csv to users seen in train/submission sets, so different statistics would fit in RAM. Size of FFM train file ~180G. RAM required ~ 80G\r\n\r\n3. I have not used hashing for features, each feature seen >=10 times would get its own ID. Submission features unseen features in train are dropped\r\n\r\n4. For numerical features, I have different strategies of dealing with them - bins, log smoothing or leaving as is, best chosen for each, based on validation. Some fields contain single feature, some bag of features. Total number of features is >5M\r\n\r\n5. I added MEAP metric into libffm, so it would stop based on it, not on logloss. Used default parameters of libffm, including k=4\r\n\r\n6. No special approach was taken to deal with leak, removing  it reduces model's precision + model learned almost all leaks naturally, so no post processing was required\r\n\r\n7. Most of the features engineering was done on 32Gb laptop, later I used 128Gb RAM  machine and for the last week used AWS\r\n\r\n\r\n8. What have not worked for me - GBT leaves as features, users clustering , MF of page_views.csv, stacking of FFM. Also, run out of time for stacking ffm->xgboost. \r\n\r\n\r\n9. List of features:\r\n\r\n> country, state, platform, county, pageDocumentCategories,\r\n> pageDocumentEntities, pageDocumentTopics, publisherCTRAdv,\r\n> publisherCTRCamp, publisherCTRDoc, countryCTRAdv, countryCTRCamp,\r\n> countryCTRDoc, stateCTRAdv, stateCTRCamp, stateCTRDoc, countyCTRAdv,\r\n> countyCTRCamp, countyCTRDoc, prevDayCTRAdv, prevDayCTRCamp,\r\n> prevDayCTRDoc, currentDayCTRAdv, currentDayCTRCamp, currentDayCTRDoc,\r\n> nextDayCTRAdv, nextDayCTRCamp, nextDayCTRDoc, currentHourCTRAdv,\r\n> currentHourCTRCamp, currentHourCTRDoc, prevHourCTRAdv,\r\n> prevHourCTRCamp, prevHourCTRDoc, nextHourCTRAdv, nextHourCTRCamp,\r\n> nextHourCTRDoc, publisherSourceCTRAdv, publisherSourceCTRCamp,\r\n> publisherSourceCTRDoc, adId, documentId, campaignId, advertiserId,\r\n> userPageViewsMeta1, userPageViewsMeta2, documentIdViews,\r\n> campaignIdViews, advertiserIdViews, userDocsSeenFromLogToday,\r\n> userPageMeta1SeenToday, userPageMeta2SeenToday,\r\n> userCampSeenFromLogToday, userAdvertisersSeenFromLogToday,\r\n> userDocsSeenFromLogTomorrow, userPageMeta1SeenTomorrow,\r\n> userPageMeta2SeenTomorrow, docStats, documentCategoriesId, advStats,\r\n> documentEntitiesId, documentTopicsId, metaSourceId, metaPublisherId,\r\n> pageDocumentMeta1, pageDocumentMeta2, userClickedThisDocumentTimes,\r\n> userSkippedThisDocumentTimes, userClickedThisCampaignTimes,\r\n> userSkippedThisCampaignTimes, userClickedThisAdvertiserTimes,\r\n> userSkippedThisAdvertiserTimes, hourOfDay, dayNum, dayOfWeek,\r\n> thisAdEntityClickedBefore, daysFromAdDocPublished,\r\n> thisAdEntityClickedBeforeCount, thisAdCategoryClickedBefore,\r\n> thisAdCategoryClickedBeforeCount, userClickedThisAdTimes,\r\n> userSkippedThisAdTimes, userPageViewsDocuments,\r\n> userPageViewsCampaigns, userPageViewsAdvertisers,\r\n> userPageViewsCategories, userPageViewsEntitites, userPageViewsTopics,\r\n> seenThisCategoryInLog, userSeenThisMeta1, userSeenThisMeta2,\r\n> seenThisDocInLog, seenThisCampaignInLog, seenThisAdvInLog, campStats,\r\n> adStats, userCampSeenFromLogTomorrow,\r\n> userAdvertisersSeenFromLogTomorrow, seenThisEntityInLog,\r\n> userAllAdsFreq, nextDocUserClicked, nextCampUserClicked,\r\n> nextAdvUserClicked, userClickedCount, userSkippedCount,\r\n> userClickedCountToday, userDocsClickedToday, userCampClickedToday,\r\n> userAdvertisersClickedToday, seenThisTopicInLog,\r\n> userClickedThisDocFreq, userClickedThisCampFreq,\r\n> userClickedThisAdvFreq, lastPageViewUserSeen, nextPageViewUserSeen,\r\n> lastDocUserClicked, lastCampUserClicked, lastAdvUserClicked\r\n\r\n\r\nThanks,\r\n\r\nAndrii",
      "votes": null
    },
    {
      "id": "157194",
      "postDate": "01/19/2017 16:24:26",
      "content": "<p>@Andrii Cherednychenko  Thank you for sharing. <em>countryCTRDoc</em> means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks </p>",
      "rawMarkdown": "Andrii Cherednychenko  Thank you for sharing. *countryCTRDoc* means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks",
      "votes": null
    },
    {
      "id": "157195",
      "postDate": "01/19/2017 16:24:27",
      "content": "<p>Reviewing what is written and everything seems so easy, but devil is in the details, lot of it :-)</p>",
      "rawMarkdown": "Reviewing what is written and everything seems so easy, but devil is in the details, lot of it :-)",
      "votes": null
    },
    {
      "id": "157196",
      "postDate": "01/19/2017 16:26:21",
      "content": "<p>[quote=FengLi;157194]</p>\n\n<p>@Andrii Cherednychenko  Thank you for sharing. <em>countryCTRDoc</em> means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks </p>\n\n<p>[/quote]</p>\n\n<p>Yes, formula is correct - clicked / seen. Fighting overfit by taking this feature only when times seen is &gt; 10. Same for all CTR* features</p>",
      "rawMarkdown": "[quote=FengLi;157194]\r\n\r\n@Andrii Cherednychenko  Thank you for sharing. *countryCTRDoc* means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks \r\n\r\n[/quote]\r\n\r\nYes, formula is correct - clicked / seen. Fighting overfit by taking this feature only when times seen is > 10. Same for all CTR* features",
      "votes": null
    },
    {
      "id": "158827",
      "postDate": "01/30/2017 10:18:12",
      "content": "<p>Thanks for sharing so many rich details!</p>\n\n<p>How long did it take to execute your feature engineering and modeling on the final submission and what  instance type did you run?</p>",
      "rawMarkdown": "Thanks for sharing so many rich details!\r\n\r\nHow long did it take to execute your feature engineering and modeling on the final submission and what  instance type did you run?",
      "votes": null
    },
    {
      "id": "160771",
      "postDate": "02/09/2017 12:50:02",
      "content": "<p>Hey Andrew, sorry for delay with reply.</p>\n\n<p>Time required for feature engineering depends on number of cores, as it is designed to work in parallel, on 6th core machine it would take around 2-3 hours. Amount of data produced &gt; 500G. fast SSD drive used locally and EBS on AWS</p>\n\n<p>Same stuff with training, on 6th core machine, for k=4 it would take maybe 15 hours to train. With x8 large it would definitely take less. </p>\n\n<p>I've been using mostly x4 and x8 RAM optimized instances. Usually 1, rarely 2 in parallel. +local 128G machine (i7-6800) .</p>\n\n<p>Nearly 75% of engineering was done on 32G Alienware laptop however.</p>\n\n<p>Andrii</p>",
      "rawMarkdown": "Hey Andrew, sorry for delay with reply.\n\nTime required for feature engineering depends on number of cores, as it is designed to work in parallel, on 6th core machine it would take around 2-3 hours. Amount of data produced > 500G. fast SSD drive used locally and EBS on AWS\n\nSame stuff with training, on 6th core machine, for k=4 it would take maybe 15 hours to train. With x8 large it would definitely take less. \n\nI've been using mostly x4 and x8 RAM optimized instances. Usually 1, rarely 2 in parallel. +local 128G machine (i7-6800) .\n\nNearly 75% of engineering was done on 32G Alienware laptop however.\n\nAndrii",
      "votes": null
    },
    {
      "id": "161316",
      "postDate": "02/13/2017 05:53:02",
      "content": "<p>Great reference info! Thanks again!</p>",
      "rawMarkdown": "Great reference info! Thanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 157194,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "01/19/2017 16:24:26",
      "content": "<p>@Andrii Cherednychenko  Thank you for sharing. <em>countryCTRDoc</em> means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 157196,
          "author_name": "cherednychenko",
          "author_url": "",
          "post_date": "01/19/2017 16:26:21",
          "content": "<p>[quote=FengLi;157194]</p>\n\n<p>@Andrii Cherednychenko  Thank you for sharing. <em>countryCTRDoc</em> means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks </p>\n\n<p>[/quote]</p>\n\n<p>Yes, formula is correct - clicked / seen. Fighting overfit by taking this feature only when times seen is &gt; 10. Same for all CTR* features</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157195,
      "author_name": "cherednychenko",
      "author_url": "",
      "post_date": "01/19/2017 16:24:27",
      "content": "<p>Reviewing what is written and everything seems so easy, but devil is in the details, lot of it :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 158827,
      "author_name": "madmanminkin01",
      "author_url": "",
      "post_date": "01/30/2017 10:18:12",
      "content": "<p>Thanks for sharing so many rich details!</p>\n\n<p>How long did it take to execute your feature engineering and modeling on the final submission and what  instance type did you run?</p>",
      "votes": null,
      "replies": [
        {
          "id": 160771,
          "author_name": "cherednychenko",
          "author_url": "",
          "post_date": "02/09/2017 12:50:02",
          "content": "<p>Hey Andrew, sorry for delay with reply.</p>\n\n<p>Time required for feature engineering depends on number of cores, as it is designed to work in parallel, on 6th core machine it would take around 2-3 hours. Amount of data produced &gt; 500G. fast SSD drive used locally and EBS on AWS</p>\n\n<p>Same stuff with training, on 6th core machine, for k=4 it would take maybe 15 hours to train. With x8 large it would definitely take less. </p>\n\n<p>I've been using mostly x4 and x8 RAM optimized instances. Usually 1, rarely 2 in parallel. +local 128G machine (i7-6800) .</p>\n\n<p>Nearly 75% of engineering was done on 32G Alienware laptop however.</p>\n\n<p>Andrii</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 161316,
          "author_name": "madmanminkin01",
          "author_url": "",
          "post_date": "02/13/2017 05:53:02",
          "content": "<p>Great reference info! Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "157193": "Hey there, \r\n\r\nCongrats to the winners and thanks a lot for congrats to my side :-)\r\n\r\n\r\nSharing details on my solution if anyone is interested, basically it is FFM model averaged for last 5 runs:\r\n\r\nTechnologies / frameworks:\r\n\r\n1. Early merging of files was done using Python. Basically, I merged all files into  train/test/sumission files, then was reading line by line for generating features\r\n\r\n2. All feature generation and statistics building is done using Scala. I found it way faster than Python + I am better with it, especially when it came to parallelization. \r\nI have also filtered page_views.csv to users seen in train/submission sets, so different statistics would fit in RAM. Size of FFM train file ~180G. RAM required ~ 80G\r\n\r\n3. I have not used hashing for features, each feature seen >=10 times would get its own ID. Submission features unseen features in train are dropped\r\n\r\n4. For numerical features, I have different strategies of dealing with them - bins, log smoothing or leaving as is, best chosen for each, based on validation. Some fields contain single feature, some bag of features. Total number of features is >5M\r\n\r\n5. I added MEAP metric into libffm, so it would stop based on it, not on logloss. Used default parameters of libffm, including k=4\r\n\r\n6. No special approach was taken to deal with leak, removing  it reduces model's precision + model learned almost all leaks naturally, so no post processing was required\r\n\r\n7. Most of the features engineering was done on 32Gb laptop, later I used 128Gb RAM  machine and for the last week used AWS\r\n\r\n\r\n8. What have not worked for me - GBT leaves as features, users clustering , MF of page_views.csv, stacking of FFM. Also, run out of time for stacking ffm->xgboost. \r\n\r\n\r\n9. List of features:\r\n\r\n> country, state, platform, county, pageDocumentCategories,\r\n> pageDocumentEntities, pageDocumentTopics, publisherCTRAdv,\r\n> publisherCTRCamp, publisherCTRDoc, countryCTRAdv, countryCTRCamp,\r\n> countryCTRDoc, stateCTRAdv, stateCTRCamp, stateCTRDoc, countyCTRAdv,\r\n> countyCTRCamp, countyCTRDoc, prevDayCTRAdv, prevDayCTRCamp,\r\n> prevDayCTRDoc, currentDayCTRAdv, currentDayCTRCamp, currentDayCTRDoc,\r\n> nextDayCTRAdv, nextDayCTRCamp, nextDayCTRDoc, currentHourCTRAdv,\r\n> currentHourCTRCamp, currentHourCTRDoc, prevHourCTRAdv,\r\n> prevHourCTRCamp, prevHourCTRDoc, nextHourCTRAdv, nextHourCTRCamp,\r\n> nextHourCTRDoc, publisherSourceCTRAdv, publisherSourceCTRCamp,\r\n> publisherSourceCTRDoc, adId, documentId, campaignId, advertiserId,\r\n> userPageViewsMeta1, userPageViewsMeta2, documentIdViews,\r\n> campaignIdViews, advertiserIdViews, userDocsSeenFromLogToday,\r\n> userPageMeta1SeenToday, userPageMeta2SeenToday,\r\n> userCampSeenFromLogToday, userAdvertisersSeenFromLogToday,\r\n> userDocsSeenFromLogTomorrow, userPageMeta1SeenTomorrow,\r\n> userPageMeta2SeenTomorrow, docStats, documentCategoriesId, advStats,\r\n> documentEntitiesId, documentTopicsId, metaSourceId, metaPublisherId,\r\n> pageDocumentMeta1, pageDocumentMeta2, userClickedThisDocumentTimes,\r\n> userSkippedThisDocumentTimes, userClickedThisCampaignTimes,\r\n> userSkippedThisCampaignTimes, userClickedThisAdvertiserTimes,\r\n> userSkippedThisAdvertiserTimes, hourOfDay, dayNum, dayOfWeek,\r\n> thisAdEntityClickedBefore, daysFromAdDocPublished,\r\n> thisAdEntityClickedBeforeCount, thisAdCategoryClickedBefore,\r\n> thisAdCategoryClickedBeforeCount, userClickedThisAdTimes,\r\n> userSkippedThisAdTimes, userPageViewsDocuments,\r\n> userPageViewsCampaigns, userPageViewsAdvertisers,\r\n> userPageViewsCategories, userPageViewsEntitites, userPageViewsTopics,\r\n> seenThisCategoryInLog, userSeenThisMeta1, userSeenThisMeta2,\r\n> seenThisDocInLog, seenThisCampaignInLog, seenThisAdvInLog, campStats,\r\n> adStats, userCampSeenFromLogTomorrow,\r\n> userAdvertisersSeenFromLogTomorrow, seenThisEntityInLog,\r\n> userAllAdsFreq, nextDocUserClicked, nextCampUserClicked,\r\n> nextAdvUserClicked, userClickedCount, userSkippedCount,\r\n> userClickedCountToday, userDocsClickedToday, userCampClickedToday,\r\n> userAdvertisersClickedToday, seenThisTopicInLog,\r\n> userClickedThisDocFreq, userClickedThisCampFreq,\r\n> userClickedThisAdvFreq, lastPageViewUserSeen, nextPageViewUserSeen,\r\n> lastDocUserClicked, lastCampUserClicked, lastAdvUserClicked\r\n\r\n\r\nThanks,\r\n\r\nAndrii",
    "157194": "Andrii Cherednychenko  Thank you for sharing. *countryCTRDoc* means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks",
    "157195": "Reviewing what is written and everything seems so easy, but devil is in the details, lot of it :-)",
    "157196": "[quote=FengLi;157194]\r\n\r\n@Andrii Cherednychenko  Thank you for sharing. *countryCTRDoc* means P(clicked=1|country,document_id) ?. Then you merge it back to test data? If so, some paris are very unique which would case severe overfitting. How did you handle this issue? Thanks \r\n\r\n[/quote]\r\n\r\nYes, formula is correct - clicked / seen. Fighting overfit by taking this feature only when times seen is > 10. Same for all CTR* features",
    "158827": "Thanks for sharing so many rich details!\r\n\r\nHow long did it take to execute your feature engineering and modeling on the final submission and what  instance type did you run?",
    "160771": "Hey Andrew, sorry for delay with reply.\n\nTime required for feature engineering depends on number of cores, as it is designed to work in parallel, on 6th core machine it would take around 2-3 hours. Amount of data produced > 500G. fast SSD drive used locally and EBS on AWS\n\nSame stuff with training, on 6th core machine, for k=4 it would take maybe 15 hours to train. With x8 large it would definitely take less. \n\nI've been using mostly x4 and x8 RAM optimized instances. Usually 1, rarely 2 in parallel. +local 128G machine (i7-6800) .\n\nNearly 75% of engineering was done on 32G Alienware laptop however.\n\nAndrii",
    "161316": "Great reference info! Thanks again!"
  },
  "source": "meta"
}