BettaFish Docker Deployment Complete Guide: Build Your Own Public Opinion Analysis System
BettaFish is an innovative multi-agent public opinion analysis system that supports one-click Docker deployment. This article provides a comprehensive guide on how to quickly deploy BettaFish using Docker, including environment configuration, LLM integration, usage instructions, and troubleshooting.
Published 325 days ago. Content may be outdated.
BettaFish is quite an interesting open-source project that makes public opinion analysis as simple as having a conversation. You just need to state your analysis requirements, and the system will automatically deploy multiple AI agents to analyze content from 30+ mainstream platforms like Weibo, Xiaohongshu, TikTok, and generate a professional analysis report.
The project’s name is quite meaningful—BettaFish, just like its motto “small but powerful, fearless of challenges.” Today, let’s see how to quickly deploy this powerful public opinion analysis system using Docker.
Imagine just asking “What do people think about electric vehicles lately?” and the system analyzes discussions across the entire web for you. Pretty cool, right?
🚀 Core Advantages
- AI-Driven Comprehensive Monitoring: AI crawler clusters run 24/7 non-stop, comprehensively covering 10+ key domestic and international social media platforms including Weibo, Xiaohongshu, TikTok, Kuaishou, etc.
- Composite Analysis Engine Beyond LLM: Integrates 5 professionally designed agents and multiple middleware models
- Powerful Multimodal Capabilities: Breaks through text and image limitations, deeply analyzes short video content
- Agent “Forum” Collaboration Mechanism: Conducts ideological collision and debate through debate moderator mode
- Seamless Integration of Public and Private Domain Data: Supports seamless integration of internal business databases with public opinion data
- Lightweight and Highly Extensible Framework: Based on pure Python modular design, achieving lightweight one-click deployment
🏗️ System Architecture
Core Components
- Insight Agent: Private database mining agent
- Media Agent: Multimodal content analysis agent
- Query Agent: Precise information search agent
- Report Agent: Intelligent report generation agent
- Forum Engine: Forum collaboration engine
Complete Analysis Workflow
| Step | Phase Name | Main Operations | Participating Components | Cycle Nature |
|---|---|---|---|---|
| 1 | User Query | Flask main application receives query | Flask Main Application | - |
| 2 | Parallel Launch | Three agents start working simultaneously | Query Agent, Media Agent, Insight Agent | - |
| 3 | Preliminary Analysis | Each agent uses dedicated tools for overview search | Each Agent + Dedicated Toolsets | - |
| 4 | Strategy Formulation | Develop segmented research strategies based on preliminary results | Internal Decision Modules of Each Agent | - |
| 5-N | Iterative Phase | Forum Collaboration + In-depth Research | ForumEngine + All Agents | Multi-round cycles |
| N+1 | Result Integration | Report Agent collects all analysis results | Report Agent | - |
| N+2 | Report Generation | Dynamically select templates and styles, generate final reports through multiple rounds | Report Agent + Template Engine | - |
🐳 Docker Quick Deployment
System Requirements
- Operating System: Windows, Linux, MacOS
- Docker: Docker 20.10+ and Docker Compose
- Memory: Recommended 4GB+
- Storage Space: At least 10GB available space
1. Get Project Code
# Clone the project repository
git clone https://github.com/666ghj/BettaFish.git
cd BettaFish
2. Environment Configuration
2.1 Copy Configuration File
# Copy environment configuration file
cp .env.example .env
2.2 Edit Configuration File
Edit the .env file and configure the following key parameters:
# ====================== Database Configuration ======================
# Database host (use service name in Docker environment)
DB_HOST=db
# Database port
DB_PORT=5432
# Database username
DB_USER=bettafish
# Database password
DB_PASSWORD=bettafish
# Database name
DB_NAME=bettafish
# Database type
DB_DIALECT=postgresql
# ====================== LLM Configuration ======================
# Insight Agent LLM configuration
INSIGHT_ENGINE_API_KEY=your_api_key_here
INSIGHT_ENGINE_BASE_URL=https://api.openai.com/v1
INSIGHT_ENGINE_MODEL_NAME=gpt-3.5-turbo
# Query Agent LLM configuration
QUERY_ENGINE_API_KEY=your_api_key_here
QUERY_ENGINE_BASE_URL=https://api.openai.com/v1
QUERY_ENGINE_MODEL_NAME=gpt-3.5-turbo
# Media Agent LLM configuration
MEDIA_ENGINE_API_KEY=your_api_key_here
MEDIA_ENGINE_BASE_URL=https://api.openai.com/v1
MEDIA_ENGINE_MODEL_NAME=gpt-4-vision-preview
# Report Agent LLM configuration
REPORT_ENGINE_API_KEY=your_api_key_here
REPORT_ENGINE_BASE_URL=https://api.openai.com/v1
REPORT_ENGINE_MODEL_NAME=gpt-4
# Forum Host LLM configuration
FORUM_HOST_API_KEY=your_api_key_here
FORUM_HOST_BASE_URL=https://api.openai.com/v1
FORUM_HOST_MODEL_NAME=gpt-3.5-turbo
# ====================== Search Engine Configuration ======================
# Serper API Key (for web search)
SERPER_API_KEY=your_serper_api_key
# ====================== Other Configuration ======================
# Flask application secret key
FLASK_SECRET_KEY=your_secret_key_here
# Log level
LOG_LEVEL=INFO
3. Start Services
3.1 Start All Services in Background
# Start all services (run in background)
docker compose up -d
3.2 Check Service Status
# Check service running status
docker compose ps
# View service logs
docker compose logs -f
# View specific service logs
docker compose logs -f app
docker compose logs -f db
4. Access Application
After services start successfully, you can access through the following addresses:
- Main Application Interface: http://localhost:5000
- Database Management: PostgreSQL runs on localhost:5432
⚙️ Detailed Configuration Instructions
Database Configuration
The system supports PostgreSQL (recommended) and MySQL databases:
| Configuration Item | Docker Environment Value | Description |
|---|---|---|
DB_HOST | db | Database service name (defined in docker-compose.yml) |
DB_PORT | 5432 | PostgreSQL default port |
DB_USER | bettafish | Database username |
DB_PASSWORD | bettafish | Database password |
DB_NAME | bettafish | Database name |
| Others | Keep Default | Other parameters like database connection pool settings should be kept at default values |
LLM Model Configuration
All LLM calls in the system use the OpenAI API interface standard. You can use:
- Official OpenAI API
- Other service providers compatible with OpenAI format (such as Azure OpenAI, Claude, domestic large models, etc.)
- Locally deployed models (such as Ollama, vLLM, etc.)
Recommended LLM API service providers:
- aihubmix
- Azure OpenAI Service
- Alibaba Cloud Tongyi Qianwen
- Tencent Cloud Hunyuan
Web Search Configuration
The system uses Serper API for web search, which requires:
- Register a Serper account
- Get API Key
- Configure
SERPER_API_KEYin the.envfile
🎯 Usage Guide
Basic Usage Flow
- Access Main Interface: Open http://localhost:5000
- Input Analysis Requirements: Enter your public opinion analysis requirements in the dialog box
- Wait for Analysis Completion: The system will automatically call multiple agents for analysis
- View Analysis Report: The system will generate a detailed HTML format analysis report
Usage Examples
Example 1: Brand Public Opinion Analysis
Please analyze the online public opinion of “Huawei phones” in the past week, including user reviews, hot topics, and sentiment trends
Example 2: Hot Event Analysis
Analyze the latest discussions related to “artificial intelligence development”, focusing on technology trends and public attitudes
Example 3: Competitive Analysis
Compare and analyze the brand reputation and user feedback of “Tesla” and “BYD” in the electric vehicle field
Advanced Features
1. Using Individual Agents Separately
# Start Query Engine (search agent)
docker compose exec app streamlit run SingleEngineApp/query_engine_streamlit_app.py --server.port 8503
# Start Media Engine (multimedia analysis agent)
docker compose exec app streamlit run SingleEngineApp/media_engine_streamlit_app.py --server.port 8502
# Start Insight Engine (insight analysis agent)
docker compose exec app streamlit run SingleEngineApp/insight_engine_streamlit_app.py --server.port 8501
2. Independent Use of Crawler System
# Enter container
docker compose exec app bash
# Enter crawler directory
cd MindSpider
# Project initialization
python main.py --setup
# Run topic extraction
python main.py --broad-topic
# Run complete crawler workflow
python main.py --complete --date 2024-01-20
# Run deep crawling only
python main.py --deep-sentiment --platforms xhs dy wb
🔧 Troubleshooting
Common Issues
1. Container Startup Failure
Problem: docker compose up -d fails
Solution:
# Check Docker service status
docker --version
docker compose --version
# Clean old containers and images
docker compose down
docker system prune -f
# Rebuild and start
docker compose up -d --build
2. Database Connection Failure
Problem: Application cannot connect to database
Solution:
# Check database container status
docker compose logs db
# Check database configuration
cat .env | grep DB_
# Restart database service
docker compose restart db
3. LLM API Call Failure
Problem: Agents cannot call LLM API
Solution:
- Check if API Key is correctly configured
- Verify network connection and API service availability
- Check API quota and limits
# View application logs
docker compose logs app | grep -i error
# Test API connection
curl -H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-3.5-turbo","messages":[{"role":"user","content":"Hello"}]}' \
https://api.openai.com/v1/chat/completions
4. Port Conflict
Problem: Port 5000 is occupied
Solution:
# Find process occupying the port
netstat -tulpn | grep :5000
# Modify port mapping in docker-compose.yml
# Change "5000:5000" to "5001:5000"
# Restart services
docker compose down
docker compose up -d
5. Insufficient Memory
Problem: System runs slowly or crashes
Solution:
# Check system resource usage
docker stats
# Increase Docker memory limit
# Add in docker-compose.yml:
# deploy:
# resources:
# limits:
# memory: 4G
# Clean unused Docker resources
docker system prune -f
docker volume prune -f
Performance Optimization
1. Database Optimization
# Enter database container
docker compose exec db psql -U bettafish -d bettafish
-- Create indexes to optimize query performance
CREATE INDEX IF NOT EXISTS idx_content_created_at ON content(created_at);
CREATE INDEX IF NOT EXISTS idx_content_platform ON content(platform);
2. Application Optimization
Adjust the following parameters in the .env file:
# Reduce concurrent request count
MAX_CONCURRENT_REQUESTS=5
# Adjust timeout
REQUEST_TIMEOUT=30
# Enable caching
ENABLE_CACHE=true
CACHE_TTL=3600
📊 Monitoring and Maintenance
Log Management
# View all service logs
docker compose logs
# View logs in real-time
docker compose logs -f
# View logs for specific time range
docker compose logs --since="2024-01-01T00:00:00" --until="2024-01-02T00:00:00"
# Save logs to file
docker compose logs > bettafish.log
Data Backup
# Backup PostgreSQL database
docker compose exec db pg_dump -U bettafish bettafish > backup_$(date +%Y%m%d_%H%M%S).sql
# Restore database
docker compose exec -T db psql -U bettafish bettafish < backup_20240101_120000.sql
System Updates
# Pull latest code
git pull origin main
# Rebuild and start services
docker compose down
docker compose up -d --build
# Clean old images
docker image prune -f
🤝 Community and Support
Getting Help
- GitHub Issues: https://github.com/666ghj/BettaFish/issues
Note: This guide is written based on the current version of the BettaFish project. As the project updates, some configurations and usage methods may change. It is recommended to regularly check the project’s official documentation for the latest information.
More Articles