{
  "id": 586164,
  "title": "Project Discussion: SocialSim – Social Media Based Personas",
  "url": "/competitions/social-sim-challenge-social-media-based-personas/discussion/586164",
  "author_name": "SAIROHITH BUKKA",
  "post_date": "2025-06-25T12:08:48.391000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>🎯 Goal: Build generative and classification pipelines for persona-based social action generation.</p>\n<p>Phase: Submission phase (SOON)</p>\n<p>🔍 Problem Statement</p>\n<p>The task involves understanding and predicting social media actions (like, post, quote, reply, etc.) from multi-turn persona-based conversations. The goal is to simulate realistic social behavior using a mix of classification and generation strategies.</p>\n<p>📦 Dataset Used</p>\n<p>→Dataset provided by the SocialSim Challenge on Kaggle:</p>\n<p>→train/: Multi-label action samples (13 actions)</p>\n<p>→test/: Raw persona clusters needing response generation</p>\n<p>→persona_labels.csv: Persona tags like User, Influencer, Bot, etc.</p>\n<p>→history_text: Preprocessed multi-turn text from .json persona clusters</p>\n<p>⚙️ Modeling Strategy</p>\n<p>✅ Phase 1: Data Preprocessing</p>\n<p>→Parsed multi-turn persona clusters to extract history_text</p>\n<p>→Created final_labels_with_history.csv combining id, cluster_id, labels, and text</p>\n<p>✅ Phase 2: Baseline Classifier (TF-IDF + LogisticRegression)</p>\n<p>→Developed a lightweight classifier for quick testing</p>\n<p>→Faced data imbalance issues – resolved using selective subsampling</p>\n<p>✅ Phase 3: Sentence-BERT + MultiOutputClassifier</p>\n<p>→Used all-MiniLM-L6-v2 to embed history_text</p>\n<p>→Trained LogisticRegression-based multi-label classifier</p>\n<p>→Saved classifier_small.pkl for later action inference</p>\n<p>✅ Phase 4: Conditional Text Generator (GPT-2)</p>\n<p>→Used distilgpt2 fine-tuned locally</p>\n<p>→Conditioned generation based on action labels (label_post, label_quote, label_reply)</p>\n<p>→Introduced batching (batch size = 32) and limited to top-N active samples</p>\n<p>🌀 Challenges Faced</p>\n<p>→Huge size: 7.9M rows in test made full-generation infeasible</p>\n<p>→Class imbalance: Over 99% samples had no active actions (EMPTY)</p>\n<p>→Generation time: Model took 30+ seconds per batch without GPU</p>\n<p>→Padding issues: Resolved by setting tokenizer.pad_token = tokenizer.eos_token</p>\n<p>✅ Final Pipeline</p>\n<p>→Classification: Sentence-BERT + Logistic Regression</p>\n<p>→Conditional Generation: GPT-2 with top_k/top_p sampling</p>\n<p>→Strategic Filtering: Only generated for label_post == True rows</p>\n<p>→Partial Saving: Saved each chunk separately, merged into final CSV</p>\n<p>Feedback:<br>\nWould love to hear your thoughts! 😊<br>\nAny feedback on the model choices, generation strategy, or potential next steps is super welcome.</p>",
  "messages": [
    {
      "id": 3232139,
      "postDate": "2025-06-25T12:08:48.393Z",
      "content": "<p>🎯 Goal: Build generative and classification pipelines for persona-based social action generation.</p>\n<p>Phase: Submission phase (SOON)</p>\n<p>🔍 Problem Statement</p>\n<p>The task involves understanding and predicting social media actions (like, post, quote, reply, etc.) from multi-turn persona-based conversations. The goal is to simulate realistic social behavior using a mix of classification and generation strategies.</p>\n<p>📦 Dataset Used</p>\n<p>→Dataset provided by the SocialSim Challenge on Kaggle:</p>\n<p>→train/: Multi-label action samples (13 actions)</p>\n<p>→test/: Raw persona clusters needing response generation</p>\n<p>→persona_labels.csv: Persona tags like User, Influencer, Bot, etc.</p>\n<p>→history_text: Preprocessed multi-turn text from .json persona clusters</p>\n<p>⚙️ Modeling Strategy</p>\n<p>✅ Phase 1: Data Preprocessing</p>\n<p>→Parsed multi-turn persona clusters to extract history_text</p>\n<p>→Created final_labels_with_history.csv combining id, cluster_id, labels, and text</p>\n<p>✅ Phase 2: Baseline Classifier (TF-IDF + LogisticRegression)</p>\n<p>→Developed a lightweight classifier for quick testing</p>\n<p>→Faced data imbalance issues – resolved using selective subsampling</p>\n<p>✅ Phase 3: Sentence-BERT + MultiOutputClassifier</p>\n<p>→Used all-MiniLM-L6-v2 to embed history_text</p>\n<p>→Trained LogisticRegression-based multi-label classifier</p>\n<p>→Saved classifier_small.pkl for later action inference</p>\n<p>✅ Phase 4: Conditional Text Generator (GPT-2)</p>\n<p>→Used distilgpt2 fine-tuned locally</p>\n<p>→Conditioned generation based on action labels (label_post, label_quote, label_reply)</p>\n<p>→Introduced batching (batch size = 32) and limited to top-N active samples</p>\n<p>🌀 Challenges Faced</p>\n<p>→Huge size: 7.9M rows in test made full-generation infeasible</p>\n<p>→Class imbalance: Over 99% samples had no active actions (EMPTY)</p>\n<p>→Generation time: Model took 30+ seconds per batch without GPU</p>\n<p>→Padding issues: Resolved by setting tokenizer.pad_token = tokenizer.eos_token</p>\n<p>✅ Final Pipeline</p>\n<p>→Classification: Sentence-BERT + Logistic Regression</p>\n<p>→Conditional Generation: GPT-2 with top_k/top_p sampling</p>\n<p>→Strategic Filtering: Only generated for label_post == True rows</p>\n<p>→Partial Saving: Saved each chunk separately, merged into final CSV</p>\n<p>Feedback:<br>\nWould love to hear your thoughts! 😊<br>\nAny feedback on the model choices, generation strategy, or potential next steps is super welcome.</p>",
      "rawMarkdown": "🎯 Goal: Build generative and classification pipelines for persona-based social action generation.\n\nPhase: Submission phase (SOON)\n\n🔍 Problem Statement\n\nThe task involves understanding and predicting social media actions (like, post, quote, reply, etc.) from multi-turn persona-based conversations. The goal is to simulate realistic social behavior using a mix of classification and generation strategies.\n\n📦 Dataset Used\n\n→Dataset provided by the SocialSim Challenge on Kaggle:\n\n→train/: Multi-label action samples (13 actions)\n\n→test/: Raw persona clusters needing response generation\n\n→persona_labels.csv: Persona tags like User, Influencer, Bot, etc.\n\n→history_text: Preprocessed multi-turn text from .json persona clusters\n\n⚙️ Modeling Strategy\n\n✅ Phase 1: Data Preprocessing\n\n→Parsed multi-turn persona clusters to extract history_text\n\n→Created final_labels_with_history.csv combining id, cluster_id, labels, and text\n\n✅ Phase 2: Baseline Classifier (TF-IDF + LogisticRegression)\n\n→Developed a lightweight classifier for quick testing\n\n→Faced data imbalance issues – resolved using selective subsampling\n\n✅ Phase 3: Sentence-BERT + MultiOutputClassifier\n\n→Used all-MiniLM-L6-v2 to embed history_text\n\n→Trained LogisticRegression-based multi-label classifier\n\n→Saved classifier_small.pkl for later action inference\n\n✅ Phase 4: Conditional Text Generator (GPT-2)\n\n→Used distilgpt2 fine-tuned locally\n\n→Conditioned generation based on action labels (label_post, label_quote, label_reply)\n\n→Introduced batching (batch size = 32) and limited to top-N active samples\n\n🌀 Challenges Faced\n\n→Huge size: 7.9M rows in test made full-generation infeasible\n\n→Class imbalance: Over 99% samples had no active actions (EMPTY)\n\n→Generation time: Model took 30+ seconds per batch without GPU\n\n→Padding issues: Resolved by setting tokenizer.pad_token = tokenizer.eos_token\n\n✅ Final Pipeline\n\n→Classification: Sentence-BERT + Logistic Regression\n\n→Conditional Generation: GPT-2 with top_k/top_p sampling\n\n→Strategic Filtering: Only generated for label_post == True rows\n\n→Partial Saving: Saved each chunk separately, merged into final CSV\n\nFeedback:\nWould love to hear your thoughts! 😊\nAny feedback on the model choices, generation strategy, or potential next steps is super welcome.\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3232139": "🎯 Goal: Build generative and classification pipelines for persona-based social action generation.\n\nPhase: Submission phase (SOON)\n\n🔍 Problem Statement\n\nThe task involves understanding and predicting social media actions (like, post, quote, reply, etc.) from multi-turn persona-based conversations. The goal is to simulate realistic social behavior using a mix of classification and generation strategies.\n\n📦 Dataset Used\n\n→Dataset provided by the SocialSim Challenge on Kaggle:\n\n→train/: Multi-label action samples (13 actions)\n\n→test/: Raw persona clusters needing response generation\n\n→persona_labels.csv: Persona tags like User, Influencer, Bot, etc.\n\n→history_text: Preprocessed multi-turn text from .json persona clusters\n\n⚙️ Modeling Strategy\n\n✅ Phase 1: Data Preprocessing\n\n→Parsed multi-turn persona clusters to extract history_text\n\n→Created final_labels_with_history.csv combining id, cluster_id, labels, and text\n\n✅ Phase 2: Baseline Classifier (TF-IDF + LogisticRegression)\n\n→Developed a lightweight classifier for quick testing\n\n→Faced data imbalance issues – resolved using selective subsampling\n\n✅ Phase 3: Sentence-BERT + MultiOutputClassifier\n\n→Used all-MiniLM-L6-v2 to embed history_text\n\n→Trained LogisticRegression-based multi-label classifier\n\n→Saved classifier_small.pkl for later action inference\n\n✅ Phase 4: Conditional Text Generator (GPT-2)\n\n→Used distilgpt2 fine-tuned locally\n\n→Conditioned generation based on action labels (label_post, label_quote, label_reply)\n\n→Introduced batching (batch size = 32) and limited to top-N active samples\n\n🌀 Challenges Faced\n\n→Huge size: 7.9M rows in test made full-generation infeasible\n\n→Class imbalance: Over 99% samples had no active actions (EMPTY)\n\n→Generation time: Model took 30+ seconds per batch without GPU\n\n→Padding issues: Resolved by setting tokenizer.pad_token = tokenizer.eos_token\n\n✅ Final Pipeline\n\n→Classification: Sentence-BERT + Logistic Regression\n\n→Conditional Generation: GPT-2 with top_k/top_p sampling\n\n→Strategic Filtering: Only generated for label_post == True rows\n\n→Partial Saving: Saved each chunk separately, merged into final CSV\n\nFeedback:\nWould love to hear your thoughts! 😊\nAny feedback on the model choices, generation strategy, or potential next steps is super welcome.\n"
  }
}